conv.

All stories
SecurityActive · 21h

Researchers extract hidden reasoning from OpenAI, Anthropic, Google LLM APIs

A new technique decodes encrypted reasoning blocks from proprietary model APIs, recovering 704 privacy artifacts including API keys and passwords.

Conversation activity · last 22 hours peak 3/30m

Peak 3 items in one 30m at Aug 11, 9 AM; 33 items over 22 hours Aug 10, 11:06 PM — no itemsAug 10, 11:36 PM — 1 item · Press 1Aug 11, 12:06 AM — no itemsAug 11, 12:36 AM — no itemsAug 11, 1:06 AM — no itemsAug 11, 1:36 AM — no itemsAug 11, 2:06 AM — no itemsAug 11, 2:36 AM — no itemsAug 11, 3:06 AM — 1 item · Hacker News 1Aug 11, 3:36 AM — no itemsAug 11, 4:06 AM — no itemsAug 11, 4:36 AM — no itemsAug 11, 5:06 AM — no itemsAug 11, 5:36 AM — no itemsAug 11, 6:06 AM — no itemsAug 11, 6:36 AM — no itemsAug 11, 7:06 AM — no itemsAug 11, 7:36 AM — no itemsAug 11, 8:06 AM — no itemsAug 11, 8:36 AM — no itemsAug 11, 9:06 AM — 3 items · Press 2, Hacker News 1Aug 11, 9:36 AM — 1 item · Hacker News 1Aug 11, 10:06 AM — 1 item · Hacker News 1Aug 11, 10:36 AM — 1 item · Hacker News 1Aug 11, 11:06 AM — 4 items · Hacker News 4Aug 11, 11:36 AM — no itemsAug 11, 12:06 PM — 2 items · Hacker News 2Aug 11, 12:36 PM — 1 item · Hacker News 1Aug 11, 1:06 PM — 1 item · Mastodon 1Aug 11, 1:36 PM — 1 item · Mastodon 1Aug 11, 2:06 PM — no itemsAug 11, 2:36 PM — no itemsAug 11, 3:06 PM — 3 items · Hacker News 2, Mastodon 1Aug 11, 3:36 PM — 2 items · Hacker News 2Aug 11, 4:06 PM — 2 items · Hacker News 2Aug 11, 4:36 PM — 1 item · Hacker News 1Aug 11, 5:06 PM — 2 items · Techmeme 2Aug 11, 5:36 PM — no itemsAug 11, 6:06 PM — 1 item · Hacker News 1Aug 11, 6:36 PM — 2 items · Mastodon 1, Press 1Aug 11, 7:06 PM — 3 items · Hacker News 2, Mastodon 1Aug 11, 7:36 PM — no itemsAug 11, 8:06 PM — no itemsAug 11, 8:36 PM — no items 3 items · 9:06 AM
Aug 114 AM8 AM12 PM4 PMnow · 9:06 PM

Summary, timeline and people extracted by Claude from 33 items across 4 sources · 6h ago. Quotes are verbatim.

Researchers led by Alexander Panfilov have published a method for extracting and decoding the hidden reasoning traces that proprietary LLM APIs return in encrypted blocks. By replaying these blocks in weaker jailbroken models from the same provider, they recovered 315,320 reasoning blocks from public repositories and discovered 704 privacy artifacts, including 62 API keys, 33 passwords, and sensitive user information. The work demonstrates a significant vulnerability affecting frontier models from OpenAI, Anthropic, and Google.

  • Researchers discovered encrypted reasoning blocks from proprietary LLM APIs can be extracted and decoded by replaying them in weaker jailbroken models from the same provider.
  • The vulnerability affects frontier models from OpenAI, Anthropic, and Google, with 315,320 reasoning blocks recovered from public repositories.
  • The extracted reasoning blocks revealed 704 privacy artifacts including API keys, passwords, access tokens, and personal identifiers, with 64 artifacts appearing only in hidden reasoning and not in visible session data.
  • The decoded reasoning output closely correlates with hidden thinking token counts reported by APIs, enabling accurate reconstruction of model reasoning processes.

How it unfolded

  1. Reaction Social media amplification

    The findings are shared on Twitter by kotekjedi_ml and spread across multiple platforms including Mastodon.

  2. The research gains traction on Hacker News with 206 points and 78 comments as the technical community discusses the security implications.

  3. Event Paper appears on arXiv

    The research is posted to arxiv.org, making the full academic paper publicly available.

  4. Panfilov et al. publish findings showing how encrypted reasoning blocks from proprietary LLM APIs can be extracted and decoded, revealing hidden reasoning traces and sensitive information.

What people are saying verbatim

“Model providers return a model's reasoning to the client as an encrypted block, which is sent back to the server when the conversation continues. These blocks are portable: they can be replayed outside their original context.”

Alexander Panfilov et al., Research authors · stolen-thoughts.com ↗

“Injecting one into a weaker, jailbroken model from the same provider allows us to extract the stronger model's raw reasoning verbatim.”

Alexander Panfilov et al., Research authors · stolen-thoughts.com ↗

“These hidden traces contain real secrets and sensitive information. Restricting to genuine, non-benchmark user sessions, we recovered 704 distinct privacy artifacts, including 62 API keys, 33 passwords, 24 access tokens, and 30 personal email addresses.”

Alexander Panfilov et al., Research authors · stolen-thoughts.com ↗

“Of those 704 artifacts, 64 appeared exclusively inside the reasoning blocks and nowhere in the visible session.”

Alexander Panfilov et al., Research authors · stolen-thoughts.com ↗

Voices from the web unedited

  • I am slowly turning around on the idea of opaque reasoning tokens.In principle, yes, I want total control and visibility into the reasoning process.In practice, I find that it takes up so much time to DIY reasoning agents that I can't spend much energy on the actual application.The model providers have way more resources and talent to do this…

    bob1029Hacker News4h agoview on Hacker News ↗
  • The authors point out several reasons why this is convenient for both the providers and users, but even without the discovery that they could prompt lower end models to spit out the contents, it seems like using one global secret should have set off alarm bells

    reedmideke@mastodon.socialMastodon · toot.community1h agoview on Mastodon ↗
  • >We take a trace produced by a frontier model, replay it into a weaker sibling, jailbreak the weaker model, ...Ha! I've been wondering if replaying across models would work, ever since https://blog.cryptographyengineering.com/2026/05/29/fooling-...I'm honestly rather curious if this was intentionally allowed, it's the sort of validation that's…

    GroxxHacker News10h agoview on Hacker News ↗
  • "Stealing" something you already paid for (tokens), but that you can't have access to(!). And trained on the sum of human knowledge.Training on other model outputs ought to be business as usual, stop using morally charged terms made up by future monopolists:

    AissenHacker News5h agoview on Hacker News ↗
  • Apparently you can do the same by simply running it without reasoning, while giving it a thinking tool...>guys you do know you can just disable thinking, and instead give it a "deep_think" tool, and it will call it with internal CoT reasoning format right?>gl fixing that

    PragmataHacker News5h agoview on Hacker News ↗
  • > For some AIME problems Opus 4.8 sometimes states the answer before deriving it. We find that the API summary does not always preserve this distinction, and can instead make the reasoning appear like a clean derivation.No surprise here but good to have more confirmation that they just put all that in the training data. And based on the…

    vhantzHacker News8h agoview on Hacker News ↗
  • I did this with Codex's recent encryption of compaction.Interestingly, I didn't have to drop to a dumber model, just a 2 sentence <developer> prompt auto-injected before and after compaction made all their models output the encrypted compaction data in plaintext.The result was... interesting. There's nothing unique in there and I still don't…

    glubHacker News1h agoview on Hacker News ↗
  • Is this how the eastern labs "distill" SOTA models?If you can play it right, you don't even need to send suspicious prompts to the frontier models. Just use them for regular tasks, extract the encrypted COT blocks and replay it to a cheaper model to get the plain text COT.But the real question is: Is it okay to steal from a thief's hoard?

    myworkaccount2Hacker News9h agoview on Hacker News ↗
  • > stop using morally charged terms made up by future monopolistsLets not gloss over this claim. Being: “Stealing is a morally charged term made up by future monopolists.”I strongly disagree. Stealing is not a made up term and property rights are foundational for any society. Your take is at least sensationalist if not malicious.

    nonethewiserHacker News2h agoview on Hacker News ↗
  • "Stealing" is a strong word to use for looking at the words produced by models built from the collective commons of the world.And, honestly, being able to see how LLMs make decisions is critical to trust and security. I consider it a valuable feature, somewhat akin to seeing the source of software I use.

    SwellJoeHacker News9h agoview on Hacker News ↗