Closed models are also used for nefarious usage.
You could limit what it was training on in the first place - however that would damage capability - and it's difficult to curate the input, especially when the models can do 2+2. ie the choice is between model power and safety - and they choose power and everything else is a sticking plaster.
One thing I find amusing is the refusal of a lot of the models to now output a lab based protocol because of fears about 'weapons' - yet I can buy a textbook or simply read papers for exact protocols.
I find it hard to reason that a person who isn't motivated enough to read a paper or buy a book, is somehow enabled to make a biological weapon because of ChatGPT - despite them needed to buy a whole bunch of specialist equipment and reagents to do it.
Are there a whole bunch of proto-terrorists who are frustrated simply because they don't know where to start?
Maybe the only place their might be radicalized teenagers - but then that's perhaps a reason for keeping them off the internet full stop :-)
The knowledge to create weapons is already widespread, the idea that terrorists need chatgpt for that is laughable
No surprise here but good to have more confirmation that they just put all that in the training data. And based on the "reasoning", the models have some form of index of those problems (or they are HEAVILY trained on them).
The real question is whether or not the training was directed to optimize for those benchmarks.
The technique doesn't guarantee that the reasoning is returned verbatim because it relies on the weaker model transcribing it accurately. Looking at the charts, there are a lot of dots that aren't in the 1:1 line that suggests that the output is exactly what was provided.
At least in the EU there is no copyright for LLM outputs, so I guess all they might do is violate the terms of service.
Language is all just sounds and markings. Anything can be redefined to mean anything, and anyone can decide to aggressively assert their preferred definition of a word.
Of course, you can assert that the meaning of "steal" only applies to physical items. You are well within your right to do so. You'd be wrong, but you can do it.
To be fair when someone tries to shift the meaning of words everyone doesn't just have to go with it to appease the large corporations trying to do that. I of course don't mean intellectual property rights or copyright infringement, you can perhaps apply the word "steal" there, not when talking about LLM traces which are currently legally uncopyrightable, though. Unless we're actually talking about someone breaking into Anthropic's servers and stealing their files, then again... if you do that you can always just blame the LLM you used.
This is the same use as "the baseball player stole third base". Nobody is depriving anyone of anything, nobody is committing a crime. It is simply: someone has obtained something in a way someone else did not intend.
There's no legal claim being made here, you have made it up.
It's more than that. By claiming that copyright infringement isn't stealing, they're usually doing so to justify such behavior: if the original thing remains with the owner, it couldn't have harmed him, could it?
(And yes, "legally" matters, because we're talking about laws in this thread, not colloquial "their life was stolen" type expressions.)
This question is obviously (hopefully) rhetorical, no need to answer. My point is that different crimes are different. Otherwise literally every crime is stealing, and no other words for different crimes matter. Obviously different crimes are different.
In most U.S. states, the actual crime will be a specific reference to a section in a Penal Code (or, for Federal crimes, the U.S. Code). For civil actions, it's likely to be a reference to a common-law tort, or some Federal statute providing a private right of civil action.
In the case of taking a physical object from someone else, most states call it "theft" in the penal code, or "conversion" for the common-law tort.
But all of this is academic anyway. I'm not entirely sure what your point is.
But to answer your question more directly, here's the most common example: https://en.wikipedia.org/wiki/Theft_of_services
And another for good measure: https://en.wikipedia.org/wiki/Identity_theft
The real issue with Anthropic, OpenAI etc. is not that they have used all of our public knowledge for training their LLMs. Creating new work from old and learning from prior generations is what we all do. The issue is that they want to claim all of the benefits for themselves. They are standing on the shoulders of giants and have contributed an inch themselves, yet want to privatize the power of the whole giant. We shouldn't let them "own" these models.
The influence on society by AI is so novel that it's reasonable to craft new laws specifically for them. There are a lot of ways to deal with their power grab. We could force them to open source the models after two years. Or we could tax tokens or compute. We just need to agree that the power grab is the problem, the privatization of our cumulative knowledge, and not some details about copyright infringement.
[1] I know I'm going to risk dissent just by putting quotation marks here. But I think for this topic specifically it is crucial to understand that intellectual property is an arbitrary social/legal construct. With physical stuff, there is an inherent scarcity. If you steal my smartphone, I no longer have it. If you steal the character from my book, I... have a harder time selling my next book? Our ancestors have invented copyright to solve a specific problem, but the solution has become perverted over time. There are a lot of egregious cases out there (looking at you, Disney), but even relatively tame success cases don't look good. Society has paid J.K. Rowling a literal billion for her work and still this cultural touchstone of a generation remains privatized. Imagine what other authors could have build upon her stories, if only they were allowed to publish their own stories with these characters. She has not been a particularly good steward in the past decades.
No, otherwise there would be a straightforward "non-commercial" clause. Instead there's a 4 part test, which takes usage (commercial or not) into account, but doesn't hinge solely on it.
>... In determining whether the use made of a work in any particular case is a fair use the factors to be considered shall include:
>1. the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes;
>...
If they really meant "non-commercial use only", they sure did spend a lot of words to not say that.
For example, consider my browser cookies that authenticate me to HN right now. Nobody even wants to copyright them, but if you were to somehow acquire a copy I'd very much consider it "stealing."
And, honestly, being able to see how LLMs make decisions is critical to trust and security. I consider it a valuable feature, somewhat akin to seeing the source of software I use.
>guys you do know you can just disable thinking, and instead give it a "deep_think" tool, and it will call it with internal CoT reasoning format right?
>gl fixing that
Interestingly, I didn't have to drop to a dumber model, just a 2 sentence <developer> prompt auto-injected before and after compaction made all their models output the encrypted compaction data in plaintext.
The result was... interesting. There's nothing unique in there and I still don't understand why they decided to encrypt it in the first place.
There's nothing to gain from this, really. Perhaps they're preparing for something in the future, where they could give the model server-side tools that improves summarization, but right now, it's just a simple prompt.
Training on other model outputs ought to be business as usual, stop using morally charged terms made up by future monopolists: https://thomasdullien.github.io/posts/2026-06-15-rl-economic...
Anyway, you can distinguish this from the debate over copyright.
They explicitly do not promise reasoning traces. You (general you) agree to those terms and pay for that bargain anyways.
But personally it’s not about right and won’t it’s just blatant bullshit.
Just because you paid for the lawyer time/LLM tokens doesn't mean you get access to everything that happened within that time/tokens.
> but you’re charged for the report itself separately from the summary, and you are not allowed to access the unsummarized report
This analogy works if the LLM provider promises you access to the reasoning tokens, and fails if they don’t.
Whether that violates the ToS is another matter Anthropic is of course free to sue for damages or stop doing business with you.
Alternatively, I paid for the tokens therefore I should have access to them. If the vendor wants to artificially hide them from me, I'll just find another way to access them.
No, I literally am paying for the thought process, per token. "Pay only for the result" is not how these things are billed.
If I was being charged for the raw, output/input token count, excluding thinking/reasoning token costs, then sure. But at least via the API, you pay for tokens you cannot see.
There are features of input and output that are opaque to you, but that you pay for. Part of how model providers chose to run their service.
Lets not gloss over this claim. Being: “Stealing is a morally charged term made up by future monopolists.”
I strongly disagree. Stealing is not a made up term and property rights are foundational for any society. Your take is at least sensationalist if not malicious.
If you broke into a data center, pulled a hard drive and drive off with it - that’s stealing. If you accessed a copy of some information - that’s infringement, unauthorized access, or some other violation. But that’s not “stealing”, which fundamentally requires a loss or otherwise depriving original owner of the property that was stolen.
Do you think stealing is real?
Compare this to “the smell of soup and the sound of money” type “theft”.
I don't "think" it's not real. I know.
Do you agree with what he actually said?
The current copyright status quo has them placed in the public domain. There is literally nothing wrong with "stealing" those tokens. They have exactly zero legal protection. "Stealing" those AI output tokens is so fundamentally impossible that it wouldn't be "stealing" in this case even if you subscribe to the copyright monopolist propaganda that copyright infringement is "stealing", and I most certainly do not.
Hilariously, that means we don't even fall prey to things like DMCA anticircumvention laws. If they encrypt the reasoning traces and we break the encryption somehow, we've done nothing wrong since the data wasn't copyrighted in the first place!
This thread and this entire topic isn’t about stealing physical goods or money. We can all agree that if I break into your house and take your TV then that’s the ancient, obvious crime of stealing.
Grice’s maxims and common sense indicate that we’re talking about the word “stealing” as applied to infringement or unauthorized copying.
arguable, and even more tenuous for intellectual "property", which was a relatively recent invention. plenty of interesting arguments over this way back to even the 19th century.
- This is about me
- I created this
- Neither of the above, but according to some story I get to control who sees it
Maybe some of those ideas are worth building into our society, but let's not pretend that The Code of Hammurabi gave a damn about intellectual property. IP was invented by the church so they could censor editions of the bible they didn't like and has been used to justify similar kinds of censorship ever since.
The foundational thing about property is that when it gets stolen, somebody else has it, and you don't.
We can hold the AI companies responsible for their actions without contributing to notions about property that encourage censorship.
Stealing pertains to more than just data. I think you agree with me that stealing isnt a made up term.
Please reread what he said. He didnt say “distilling” was morally charged made up term by future monopolists. He said “stealing.” Thats insane.
You're fixating on a few clumsily placed words and coming away with a meaning which that poster did not intend. Consider absorbing the whole context before going on the offensive. The link they shared makes it pretty clear what they were trying to say even if they fumbled the words a little.
Is this stealing? Is it depriving NYT or publishers/writers from money via lost sales/subs? I don't know, but it certainly could be.
[0] https://www.nytimes.com/2023/12/27/business/media/new-york-t...
Mass downloading copyrighted works is. Which they did. Aaron got threatened with 20 years, they got pentagon contracts.
Dont you agree thats either sensationalist hyperbole or a genuinely crazy idea?
If you can play it right, you don't even need to send suspicious prompts to the frontier models. Just use them for regular tasks, extract the encrypted COT blocks and replay it to a cheaper model to get the plain text COT.
But the real question is: Is it okay to steal from a thief's hoard?
By definition it cannot be stealing since you're paying for the tokens. It may be against their ToS, depending on what you end up doing with those tokens, but it cannot be stealing. If they charge by the token, all your tokens are belong to you :)
I also find it very strange that everyone sort of accepts their ToS like no big deal. Imagine MS using the same terms for their software - you cannot use any MS software to develop competing services. Bananas! They'd be dragged through the courts like it's the 90s.
(I get why they're doing it. Distillation is unreasonably effective. But still, I find it bananas that we've kinda accepted it, to the point where people use "stealing" or "attack" or any such terms)
(though maybe there's another interpretation of the thought alignment?)
I’m not sure this argument is correct. You can sign whatever contract you like with the model provider, right? Including “you are entitled to the end product but not the intermediate scratch work”?
Coming from a place of genuine curiosity: is there some precedent or statute that would invalidate that contract? I don’t see why the reasoning tokens belong to you.
For example, I pay lawyers by the hour but don’t necessarily own their meeting minutes, recorded discussions, research notes, etc.
Are you a lawyer?
Sure you can set spending limits, just like you can make an account and give it a limited amount of credits.
How does this relate to your previous paragraphs? LLM outputs are not copyrightable and you didn't break into Anthropic servers to steal the files from there. So how exactly is it theft? If I send an "encrypted" files to thousands of peoples and some manage to figure out how to read it I can't really accuse them of that or can I?
Besides that, the capabilities of a model are heavily dependent on the unsupervised learning phase, that gobbles all kind of other people's IP without giving a fuck. All the underpaid work behind the masses of third world programmers creating those post-training datasets would be completely uselless without it.
Also, it is kind of funny that labs resort to the "Research, time, money and expertise" argumet, when it is basically the same argument from publishers and other IP creator that the labs spent millions of dollars of lawyering money to resist. Besides, US law rejects in: Effort and cost by themselves not necessarely generate protectable interests.
About encryption, I think we're all contaminated by the bad ideology behind DMCA. While encryption established the intent, it doesn't follow that they have a legal claim of exclusivity just because of it.
Technically, you're overstating the value of so called "reasoning traces". You can't infer the verifier design, the reward shaping,or the data pipeline from them. Also, what you can extract are not the traces themselves, but the written summary of it, and you can't even guarantee that this summary reflects the exactly reasoning trace, models have show to have lied about it. Besides, distillation works when the student model already has strong priors, you can't turn a weak model in a SOTA with it. Don't believe Amodei's outrageous lies about it, he is just trying to exercise some regulatory capture.
I wonder if you can use these for attacks, like this previous paper showing that if you know how a model reasons, you can "fake its thinking" to control it? https://news.ycombinator.com/item?id=48631888
I assume switching the model in the middle of a conversation is intended behavior (very useful in coding agents, for example)
(Thanks for the link. That’s an interesting idea!)
I'm building agent to control machines at work and have to rely on small models, which are especially bad at remembering rules like this, so it is very apparent to me that soft rules like this are pretty much useless. Might be different for these huge models, but still it seems very shaky to me.
A harder to defend against approach here where they work backwards from the results and ask the model to generate a plausible trace: How to Steal Reasoning Without Reasoning Traces https://arxiv.org/pdf/2603.07267
They also talk about successful distillation of black box model capabilities with the approach.
This is the kind of stuff I point to when people talk about AI slop. AI is just a tool. You're still the person who has to deliver the output and have some taste.
Yes, some of them do do that. For example Moonshot tried to reward shorter reasoning traces in between Kimi-K2.6 and Kimi-K2.7 Code, and the latter has a mild caveman accent in its reasoning traces that the former lacks.
Qwen3.8-Max also has terse reasoning, but I don't remember this being the case for Qwen3.6 models I ran locally.
But very interesting result.
I was experimenting a bit how I could block these using an ingress path. GitHub /softcane/hamza
I went straight to the ‘Responsible Disclosure’ section. Not surprising, but still disappointing.
The stateless part is also important for enterprise customers that require zero data retention.
(they could scope CoT access per model, but then users couldn't switch models mid-session)
you have a secret to keep that is read by the llm, and untrusted input that wants to exfiltrate it.
by hell or high water, the agent is gonna output that text
For example, when there was a paper that came out showing that having model logprobs makes distillation an order of magnitude easier, the closed LLM providers instantly yanked out support for getting the full logprobs at every time step. You get at most top 10 candidates now and I'm sure even that's on the chopping block.
People will use this to argue that a model which has exceeded Opus 4.8 (Kimi K3) somehow got most of its performance through distillation of Opus 4.8.
I still don't buy that distillation was worth more than 3 months of "catch up" time for the chinese labs. Most people who use the word "distillation" to much are revealing their sinophobia.
I can’t fault them too much, as logit based distillation is extremely effective.
Very useful for making smaller models out of bigger open weight models.
The hiding of this data only brings distrust to their frontier models. I think most people want to understand how something comes to a conclusion they don't want have that part left out on purpose...
it's this kind of behavior that forces people move to to open source models in the end, it's the lack of trust. the frontier model providers treat the end user/customer as a threat or adversary. Fable 5 is notorious for this. a lot of the serious questions you ask the model they won't even respond to you because of the woke guardrails. it wasn't only a couple weeks ago that huggingface had to use glm 5.2 to get the right answers about their security incident because Fable 5 didn't want to answer it.
Anyway, can someone explain the part about K3? What are they trying to say?
From [1]:
> As you might guess, this suggests that distilling reasoning traces may have been possible for a long time without ever breaking the cryptography.
> An anecdote: we find that prefilling Kimi-K3 reasoning with a few tokens of Opus reasoning measurably shifts its response toward Opus’s
> A small memorization analysis showed that specific Claude and GPT reasoning spans are up to ~6 orders of magnitude easier to extract from Kimi-K3 than from the next-closest model.
[1] https://x.com/kotekjedi_ml/status/2087147042888114428?s=42
> We take a trace produced by a frontier model, replay it into a weaker sibling, jailbreak the weaker model, and recover the stronger model’s hidden reasoning in plaintext.
Should be easy for them to fix though: switch up the encryption key so it only works with the API for each specific model, rather than being shared across all of their models.
And indeed, the paper says it's been fixed by all three providers (though no news on how they fixed it):
> All model providers acknowledged the receipt of our report and subsequently we were unable to launch the same attacks.
You think that you are adding a description for your post when in fact you're simply submitting a regular comment not promoted or distinguished in any way.
It's even worse considering posts without URLs would take the same text from the same input box on the HN submission form and append it under the post title, locking it to the top of the page - so if you've browsed linkless posts you think maybe the submission text is gonna end up just like that, but it doesn't. (Also I think you'll even see URLs like an archive link end up appended just below the submission title for URL posts. Which I guess is a special feature.[1])
Any reason I should not send HN an email requesting clarification of this the submission page?
Edit: quoting https://news.ycombinator.com/submit :
“If there is no url, text will appear at the top of the thread.”
OK, now I finally understand what that means in practice, but it doesn’t imply an entirely separate undistinguished comment will be simultaneously submitted on my behalf.[1]modpowers(?) used to directly append links to URL submissions further confuse the matter: https://news.ycombinator.com/item?id=49243880
I feel bad now :(
Feedback emailed to HN!
In principle, yes, I want total control and visibility into the reasoning process.
In practice, I find that it takes up so much time to DIY reasoning agents that I can't spend much energy on the actual application.
The model providers have way more resources and talent to do this right and keep it right over time. I am willing to concede this moat to them if it means I can actually focus on the business.
The more I think about it, the less I care to see those tokens. It feels an awful lot like obsession over logging every trace item an application could produce. Useful in theory but a nasty garage stacked to the ceiling with useless shit otherwise. What nefarious things are we concerned with? That they burn too many tokens in the black box? I'll threaten to move to a different black box. There are always options in a market with this many participants and big players feel this pressure. They know there's a point at which being evil bastards is no longer profitable.
Spending time carefully designing tools and views that interact with the environment in clean ways is a much better investment than trying to own and control 100% of the reasoning process or models.
Ha! I've been wondering if replaying across models would work, ever since https://blog.cryptographyengineering.com/2026/05/29/fooling-...
I'm honestly rather curious if this was intentionally allowed, it's the sort of validation that's easy to miss (particularly if you're wading into the vibe waters). Seems like something that'd be absolutely riddled with possibilities for shenanigans.
PS Here's a conversation I had with GPT 5.6 about the paper differences. https://chatgpt.com/share/6a7b64b4-ec0c-83ea-a9d2-ab1f1a1dfe...
Wouldn’t surprise me if the providers just remove that ability and lock the model once the conversation starts.
This is only required if you want users to be able to share things with everyone and you are going for the simplest implementation.
If not you could try to keep a record of keys associated with a user, then when a new request comes in look through to see if the user has a valid key to decrypt the COT.
For explicit shares, just add the key used in that one conversation to the users valid keys. For global shares use the global keys. But that's adding more complexity to the system.
https://memory-alpha.fandom.com/wiki/B-4
LLMs briefly seemed like this too, after subscriptions made the SOTA models too cheap to meter, but before they walked back on that and introduced quotas...
1. The down side is that it cannot be used across the clients even for the same user
2. Using the same encryption key was a bad choice here, a per user key would have solved this issue for sure.
> a per user key would have solved this issue for sure
It would have helped with PII leakage, but not with plain-text trace extraction attacks, right?
The compliance rules at times are outdated and people skirt around them by following the worded rule instead of the intent.
Sucks.
The flaw is that the data is not strictly tied to user session, making the session data hijacking a lot easier.
1. Its a security issue.
2. Publicly available sessions make it much worse