Home
All posts

· AI research

From Simulating Speech to Optimizing Reasoning

Why this mattersUnderstanding what drives a model’s reasoning helps judge whether a convincing explanation reflects how it reached an answer.

For two years, "reasoning model" meant writing a long chain of English sentences before the model commits to an answer. OpenAI's o1 pioneered this, DeepSeek-R1 popularized it, and by mid-2026 every flagship model (GPT-5.6, Qwen3, Kimi K2.5, GLM-5) ships with some version of a thinking-toggle [1,2]. The implicit assumption is (CoT) text is the reasoning, Quiet-STaR even trained models to "think" at every token position independent of whether a user ever asked for visible reasoning at all [8]. I argue that this assumption is false. Three 2026 papers make a strong case against this rationale, and imply that reasoning is not based on language.

Reasoning is a location, not a sentence

The most direct result comes from a University of Virginia paper on sparse autoencoders (SAEs) applied to LLaMA-3 [3]. SAEs decompose a model's internal activations into a large bank of mostly-inactive, human-interpretable directions. The Virginia group found a small set of these directions that are causally tied to reasoning behavior. Steering one direction pushes the model to solve multi-step math correctly on its own. The reasoning-relevant state activates early in generation and overrides prompt instructions telling the model not to reason explicitly [3].

How would we find a latent configuration that triggers "CoT prompting-like reasoning"? The token-level trace is one interface to that configuration. If you've spent time with mechanistic interpretability, the capabilities that look behaviorally complex are often geometrically just a handful of directions in a high-dimensional space, weighted and summed. Turpin et al. showed CoT explanations can be systematically influenced by biasing features the model never mentions. They reorder multiple-choice options so the answer is always "(A)," and the model finds a confident-sounding rationale for (A) without ever citing the reordering [4]. Anthropic's 2025 follow-up quantified this across six kinds of hints slipped into prompts, where models revealed their actual reliance on the hint in their visible reasoning trace less than 20% of the time, and outcome-based RL improved this rate initially, then plateaued well short of full faithfulness [5,6]. The cleanest version of this point is quite funny because transformers trained on hard algorithmic tasks perform just as well when the chain-of-thought tokens are replaced with meaningless repeated filler (literally "......") as when they're replaced with genuine step-by-step text, on tasks where the benefit comes from extra sequential computation steps [7].

In other words, the CoT is not a mind reflecting on a problem but closer to a hidden system prompt holding a gun to the model's head, forcing it to type out a scratchpad so that its own next-token prediction has a statistical rail to slide down. It reads its own freshly generated chunk and goes along with it, the same way it goes along with any other text that happens to precede the current position. DeepSeek-R1's paper described a moment where the model "realizes" a mistake and self-corrects, and that phrasing is sticky because we are pattern-matching machines for minds. We're tuned by evolution to over-detect agency and interiority, because falsely seeing a mind that isn't there is cheap and missing one that is can be costly, the same bias that makes us see faces in clouds.

But "the transcript is not a faithful narration" does not imply "there is no real computation happening." The UVA result above is not a claim about text at all; it's a claim about a causal latent direction that exists whether or not any token gets printed. Anthropic's own attribution-graph work on Claude 3.5 Haiku found the model planning several tokens ahead of what it actually writes: in poetry, it identifies a target rhyme word before composing a single word of the line, then writes backward toward that target, and suppressing the planning feature measurably changes the model's word choice. The plan is causally load-bearing, not decorative [9,10]. A follow-up scaling study found this backward-planning circuit is not universal and rarely runs more than one token deep even in fairly capable models [11], so the depth of genuine latent planning looks like a scaling variable the field is still mapping, not a switch that's simply on. A recent position paper synthesizing this evidence argues the field should treat latent-state trajectories, rather than surface CoT, as the default object of study for what "reasoning" in an LLM actually is [12]. There surely is a real, causal, sometimes multi-step-ahead process worth calling reasoning, and it probably is only loosely coupled to whatever text gets printed alongside it.

Superposition instead of a single path

Standard CoT is the sequential decision process of pick a token, commit, move on. If step 4 of 20 is wrong, the rest of the chain inherits the error. UPenn and Microsoft Research's "Multiplex Thinking" paper (labelled "quantum thinking" on social media) attacks this by having the model sample several candidate next-tokens at each step and blend their embeddings into a single continuous "multiplex token," rather than collapsing immediately to one discrete choice [13]. The trick that makes this trainable is that the mechanism is self-adaptive: when the model is confident, the multiplex token is nearly discrete and behaves like ordinary CoT; when it's uncertain, the same slot compactly represents several plausible continuations without lengthening the sequence [13]. Reported gains hold from Pass@1 through Pass@1024 on math benchmarks, with shorter sequences than discrete CoT baselines [13]. The model keeps a small, weighted set of hypotheses alive in one vector instead of a decision tree in token-space. It's a genuinely different search structure worth tracking against classic tree-search-over-CoT baselines. It's also a mechanism that is structurally incapable of being "narrated" faithfully in the old sense because there is no single discrete path to report. If the reasoning-effort and faithfulness literatures above are right that the text was already a lossy proxy for the computation, Multiplex Thinking is what happens when you stop pretending the proxy needs to look sequential at all.

Compressing a plan, not a sentence

NVIDIA's Fast-ThinkAct targets a more practical real-world problem. Robots can't wait for 200 tokens of English reasoning before acting, which is solved by distilling a verbose textual chain-of-thought into a short sequence of (six, in the reported configuration) continuous latent tokens while keeping those tokens "verbalizable" (decoder can still translate them back into language for a human to audit) [14]. The measured result is up to an 89.3% reduction in inference latency versus reasoning VLA baselines, with long-horizon planning, few-shot adaptation, and failure recovery preserved [14]. The "verbalizer" constraint is interesting from a safety-engineering standpoint, keeping an inspection port open on a representation that is otherwise optimized to be as compact, and as un-language-like as possible.

Future directions

Put the three results next to each other and this pipeline emerges: (i) a classifier steers the model into a reasoning-relevant latent state, (ii) the model does a short amount of processing in a multiplexed, superposed latent form, and (iii) the result projects either to a language head or directly to an action head. This would explain why reasoning-effort scaling hits saturation for pure token-count increases faster than people expected.

Easy is boring. A scaling-law analysis published this month found that reasoning capability does not scale as predictably with parameter count as general language modeling does. The returns diminish sharply above a threshold. At this point, the architectural changes to how attention handles sequential logical steps matter more than increased compute on the current design [15]. A DeepMind paper on "reasoning collapse" found that models solving a problem correctly under one framing frequently fail on an equivalent problem stated differently. Exposing models to systematic reformulations during fine-tuning reduces this brittleness by roughly 40% on their benchmark, but doesn't eliminate it [15]. If reasoning were a truly abstract, framing-independent algebra, this brittleness shouldn't exist. This suggests that "reasoning latents" are still partly surface-pattern-bound, closer to a compressed index of training-time phrasings than a clean symbolic operator. Separately, empirical work on chain-of-thought length has repeatedly found an inverted-U relationship between trace length and accuracy, and the optimal length shrinks as models get more capable [16,17]. That's a second, independent line of evidence that brute-force verbosity was never where the reasoning was, which is exactly what the latent-steering results predict.

Interactive · Longer is not smarter

Accuracy against chain-of-thought length traces an inverted U: too short underthinks, too long overthinks. As the model gets more capable, the peak rises and its optimum moves left.

300 tok
Base
Accuracy
…
Optimal length
…

Schematic inverted-U. Independent work finds an inverted-U between trace length and accuracy, with the optimal length shrinking as models improve, evidence that verbosity was never where the reasoning lived.

Accuracy versus reasoning length as an inverted U, with a marker at the current length and a line at the optimum. reasoning length (tokens, log)

The swarm is one tired model

If text was never a faithful record of a single model's reasoning, it's worth asking what that implies for systems built by chaining several instances of a model together, "multi-agent teams" or "collaborative swarm" patterns with "horizontal" reasoning. I don't see, how this is not the same frozen model file pinged across parallel threads, each thread reading a shared, fast-growing text log and taking turns guessing the next line based on a hidden system prompt that assigns it a role. No communication (the definition matters) happens exceeding a shared transcript. I imagine a swarm like the architectural equivalent of a lonely kid playing both sides of a chessboard.

One layer down, an agent never actually decides to keep working (there is no autonomous agency) but a background script runs a hardcoded while True loop that repackages the entire conversation history and shoves it back into the model's context window at every tick. A static algorithm looks at that text file and predicts the next logical step [18]. "Tool execution" is not the model operating a system but API parameters force a string conforming to a JSON schema, which an external program parses and hands to a local interpreter to actually run [19,20]. A single missed trailing comma and the entire pipeline breaks because nothing about the apparent autonomy survives outside the string. When an agent hits a terminal error and "fixes itself", the orchestration script simply caught the operating system's stderr crash message, wrapped it as another observation, and reinjected it into context. The model again runs next-token prediction on a new piece of text that happens to describe a failure [18].

A 2026 study of agent scaling found that adding more homogeneous agents (same weights, different persona prompts) produces sharp diminishing returns, while genuinely heterogeneous configurations using different base models kept improving; two heterogeneous agents matched or beat sixteen homogeneous ones [21]. A companion study of over 10,000 multi-agent research proposals found that persona-based discussion among instances of the same model actively contracts the solution space over rounds of interaction. Dense back-and-forth acts as a "semantic black hole" (a premature consensus is reached because every persona shares the same underlying priors no matter the role) [22]. A related study on "persona collapse" found that when a model roleplays a richly specified character, it systematically discards most of the specified traits and keeps only the most stereotyped ones [23]. If Multiplex Thinking is what happens when you let one model hold several genuine hypotheses in superposition, the honest reading of the multi-agent literature is that spinning up parallel roleplay threads on that same model requires heterogeneous configurations and is comparatively expensive and error-prone.

The parallel axis: reasoning effort

2026's flagship models converged on reasoning-effort settings, with six discrete levels from Light to Ultra on GPT-5.6 by prepending an instruction to the system prompt. Reinforcement learning with verifiable rewards (RLVR) applied a different per-token cost penalty for each effort level during post-training [1,2]. DeepSeek V4 trained three separate effort specialists (Non-think, Think High, Think Max) with distinct context windows and length penalties, then distills them into one checkpoint [1]. Thinking Machines' Inkling model replaced discrete labels with a continuous effort value between 0.2 and 0.99, adjusting the per-token cost coefficient λ(e) directly inside the RL reward [1], while Nemotron 3 Ultra and Qwen3's "Thinking Mode Fusion" SFT stage exposes the model to both <think>...</think> and empty-tag examples so a downstream tokenizer flag can hard-switch reasoning on or off [1].

The useful conceptual move here, per Raschka's breakdown, is separating training-compute scaling (model choice Luna/Terra/Sol/...) from inference-time scaling (the effort a fixed model spends per query). A smaller model at high effort can match a larger model at low effort on some benchmarks [1]. This is the already-deployed version of "optimizing reasoning" without latent steering. Rewarding a model for using more tokens under a "high effort" label is not the same as rewarding it for those tokens accurately reflecting its computation. A recent proposal called VERITAS shows that rewarding faithfulness directly as part of the RL objective improves task performance [24]. Effort-conditioning and faithfulness-conditioning could eventually be trained as complementary objectives rather than the industry only optimizing one.

Interactive · Effort is not faithfulness

Dialling up "reasoning effort" buys accuracy with diminishing returns. It does not make the visible trace reflect the computation, unless faithfulness is rewarded as its own objective.

Medium
Accuracy
…
Faithfulness
…

Schematic. Effort runs Light→Ultra (six levels on GPT-5.6); accuracy saturates, while faithfulness stays near the measured <20% ceiling unless trained directly.

Two curves over reasoning effort: accuracy saturating high, faithfulness staying low unless faithfulness is rewarded. reasoning effort → accuracy faithfulness

AI for Science as the consolidation test case

AI4S stress-tests model reasoning against real-world heterogeneity. S1-Omni is a recent attempt to fold domain-specific scientific reasoning (materials CIFs, SMILES strings, protein sequences, spectra, scientific images) into one model via a shared representation space, knowledge-aligned training on scientific laws, and task-specific decoders for property prediction, spectrum-to-molecule generation, and structure prediction [25]. It seems like the unified representation (+decode) architecture generalizes across domains that don't share a token vocabulary but this is a different story.

Closing thoughts

Theoretical physicist working with current models claim strong reasoning systems won't just answer questions faster, they'll compress existing theory into simpler forms, the way a modern proof of an old theorem is often shorter than the original [26]. Chess engines didn't end human play but raised the ceiling of what human+engine can find, and grandmasters now train with the tool rather than merely being replaced by it [26]. To recount one example in the line of recent human+model advances, a team of physicists with GPT-5.2 Pro collapsed a 32-variable gluon amplitude expression into a single line and propose "the obvious generalization" for arbitrary particle count [27].

No matter if a model independently judges a proof as beautiful rather than merely correct, the faithfulness gap, latent steering, multiplexed hypotheses, and persona collapse in swarms is best read as a research program built around the question "does it win". If reasoning really is migrating from cloud to a latent operator, "which simplifications are worth making" is precisely the kind of aesthetic judgment that operator will eventually need to encode and that will have the same "beauty" as a well-thought chess move from Stockfish.

Sources

  1. Raschka, S. "Controlling Reasoning Effort in LLMs." Ahead of AI, 2026. https://magazine.sebastianraschka.com/p/controlling-reasoning-effort-in-llms
  2. Raschka, S. "GPT 5.6 Configurations and Defaults." 2026. https://sebastianraschka.com/blog/2026/gpt-5-6-configurations.html
  3. He, Z., Xiong, G., Liu, B., Sinha, S., Zhang, A. (University of Virginia). "Reasoning Beyond Chain-of-Thought: A Latent Computational Mode in Large Language Models." arXiv:2601.08058, 2026. https://arxiv.org/html/2601.08058v1
  4. Turpin, M., et al. "Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting." NeurIPS 2023. https://proceedings.neurips.cc/paper_files/paper/2023/hash/ed3fea9033a80fea1376299fa7863f4a-Abstract-Conference.html
  5. Anthropic. "Reasoning Models Don't Always Say What They Think." arXiv:2505.05410, 2025. https://arxiv.org/abs/2505.05410
  6. Marktechpost. "Chain-of-Thought May Not Be a Window into AI's Reasoning: Anthropic's New Study Reveals Hidden Gaps." 2025. https://www.marktechpost.com/2025/05/19/chain-of-thought-may-not-be-a-window-into-ais-reasoning-anthropics-new-study-reveals-hidden-gaps/
  7. Pfau, J., Merrill, W., Bowman, S. R. "Let's Think Dot by Dot: Hidden Computation in Transformer Language Models." arXiv:2404.15758, 2024. https://arxiv.org/abs/2404.15758
  8. Zelikman, E., et al. "Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking." arXiv:2403.09629. https://arxiv.org/html/2403.09629v1
  9. Anthropic. "Tracing the Thoughts of a Large Language Model." 2025. https://www.anthropic.com/research/tracing-thoughts-language-model
  10. Anthropic / Transformer Circuits. "On the Biology of a Large Language Model." 2025. https://transformer-circuits.pub/2025/attribution-graphs/biology.html
  11. "Latent Planning Emerges with Scale." arXiv:2604.12493. https://arxiv.org/html/2604.12493v1
  12. Wang, W. "LLM Reasoning Is Latent, Not the Chain of Thought." arXiv:2604.15726, 2026. https://arxiv.org/abs/2604.15726
  13. Tang, Y., Dong, L., Hao, Y., Dong, Q., Wei, F., Gu, J. (UPenn / Microsoft Research). "Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge." arXiv:2601.08808, 2026. https://arxiv.org/abs/2601.08808
  14. Huang, C.-P., Man, Y., Yu, Z., Chen, M.-H., Kautz, J., Wang, Y.-C. F., Yang, F.-E. (NVIDIA). "Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning." arXiv:2601.09708, 2026. https://arxiv.org/abs/2601.09708
  15. "AI Research Highlights: The Breakthroughs of August 2026" (summarizing the MIT/Stanford/Allen Institute CoT scaling-law analysis and the DeepMind reasoning-collapse paper). https://skycrumbs.com/blog/ai-research-august-2026
  16. Wu, Y., Wang, Y., Ye, Z., Du, T., Jegelka, S., Wang, Y. "When More is Less: Understanding Chain-of-Thought Length in LLMs." arXiv:2502.07266. https://arxiv.org/abs/2502.07266
  17. "When More Thinking Hurts: Overthinking in LLM Test-Time Compute Scaling." arXiv:2604.10739. https://arxiv.org/html/2604.10739v1
  18. Roelants, P. "Implement a Simple ReAct Agent Using OpenAI Function Calling." https://peterroelants.github.io/posts/react-openai-function-calling/
  19. Jacar. "OpenAI Function Calling: Structured Output." https://jacar.es/en/openai-function-calling-structuring-model-output/
  20. DeepWiki, ombharatiya/ai-system-design-guide. "Agent Architecture and Core Patterns." https://deepwiki.com/ombharatiya/ai-system-design-guide/5.1-agent-architecture-and-core-patterns
  21. "Understanding Agent Scaling in LLM-Based Multi-Agent Systems." arXiv:2602.03794. https://arxiv.org/html/2602.03794v1
  22. "Diversity Collapse in Multi-Agent LLM Systems: Structural Analysis of Ideation." ACL Findings 2026. https://aclanthology.org/2026.findings-acl.13.pdf
  23. "Investigating Persona Collapse and Homogenization in Large Language Models." arXiv:2604.24698. https://arxiv.org/html/2604.24698v1
  24. "Rewarding Faithful Reasoning in Retrieval-Augmented Generation" (VERITAS). arXiv:2510.13272. https://arxiv.org/html/2510.13272v1
  25. "S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation." arXiv:2607.15686. https://arxiv.org/html/2607.15686v1
  26. Brown, A. "Training Sand to Think: Artificial General Intelligence & Future of Physics." Perimeter Institute talk, 2026. https://www.youtube.com/watch?v=Mw60FH5iflI
  27. Krohn, J. "AI Making Theoretical Physics Breakthroughs." Five-Minute Friday, Super Data Science, 2026. https://www.youtube.com/watch?v=mKC3XyQj51U

What should I call you?

Choose a display name for your comments. No email or account signup.

Your commenting identity

Use at least 12 characters. You’ll need this passphrase to restore the file.