· LLMs in science
I Checked 2,294 arXiv Equations for Upside-Down Fractions
Why this mattersA copied equation can change meaning without looking wrong, so reliable scientific workflows need checks beyond fluent text.
When you copy a formula from a chat window, the clipboard gets text extracted from the rendered page, not the TeX. In that text, a fraction's denominator can come before its numerator. Paste it into a second model and the model has to guess which part goes on top.
Estimates of LLM use in papers so far are based on vocabulary in abstracts and body text. Up to 17.5% of computer science abstracts showed LLM editing by early 2024,2 at least 13.5% of 2024 biomedical abstracts,3 and a majority of biomedical papers by the end of 2025.4 I wanted numbers for the equations, so I downloaded the LaTeX source of 200 arXiv papers.
Where the flip comes from
ChatGPT and several other chat interfaces render math with KaTeX. KaTeX writes each formula into the page twice: as MathML with the TeX source in an annotation, and as a visible layer of positioned boxes. The visible layer builds fractions from the bottom up, so the denominator comes first in the page text.
Text content of both layers, KaTeX 0.16.11:
| Formula | MathML layer | Visible layer |
|---|---|---|
P=TPTP+FP | P=TP+FPTP | |
softmax(QK⊤dk) | softmax(dkQK⊤) | |
k=Aexp(−EaRT) | k=Aexp(−RTEa) |
What ends up on the clipboard depends on the browser, the app, and whether the site uses KaTeX's copy-tex
extension.1 A model that only gets the visible
layer can read P=TP+FPTP as (TP+FP)/TP. For precision someone would notice. For a less
familiar formula, the inverted version compiles and looks normal.
Sample
Categories: cs.LG, cs.CV, cond-mat.mtrl-sci and eess.SP. Two months per year for 2021, 2023, 2024, 2025 and 2026. 2021 is before ChatGPT. For each category and month I opened a random page of the monthly listing and took the first five papers with LaTeX source under 25 MB. That's 200 papers, 40 per year. One of them is a 2018 paper that appeared in a later listing, and I left it out of the year comparison.
The sources contain 6,591 display equation rows (each line of an align counts separately),
44,903 inline math spans, 5,130 \frac and 1,694 slashes. 26 papers have no display
equations.
Pattern checks on all 200 papers looked for slashes followed by a product (a/bc), textbook formulas with a fixed orientation (attention scaling, precision and recall, IoU, cosine similarity, Boltzmann factors, z-scores, Adam, Bayes, Coulomb and a few others), and symbols defined as a/b in one place and b/a in another.
For the full review I took 60 papers, three per category and year, and had Claude go through every display equation with a fixed protocol. For each row it decided whether the equation, used as printed, gives something other than what the authors mean, and recorded the error type, the corrected form and the evidence in the paper. A separate pass then tried to reject each of the 189 flagged errors. 179 were kept, 3 reclassified and 7 dropped. I checked the inversions and some of the other errors by hand.
A second independent review of four papers (120 rows) found 15 rows with meaning-changing errors. 14 were the same rows as in the first review. Both reviews used the same model, so errors that model systematically misses wouldn't show up in this comparison.
I graded errors as meaning-changing (wrong sign, missing factor, wrong index) or cosmetic (visibly wrong but with only one possible reading, such as an extra closing parenthesis). The numbers below are for meaning-changing errors unless noted.
How many equations are wrong
112 of 2,294 reviewed display rows have a meaning-changing error. That's 4.9% (95% interval 3.5 to 6.4%, bootstrapped over papers). With cosmetic errors it's 7.6%.
30 of the 60 papers have none. 11 papers have errors in more than 10% of their display rows. One paper's only display equation is wrong, and another has 5 of 9 wrong.
Error types
| Failure mode | Errors | Share | Typical case |
|---|---|---|---|
| Missing or extra term or factor | 37 | 31% | a normalising 1/L that isn't there |
| Wrong index or variable | 34 | 29% | current state where the next state belongs |
| Contradicts the paper's own definition | 17 | 14% | a symbol whose definition changes halfway through |
| Other wrong operator | 15 | 13% | sup for inf, min for max |
| Sign | 9 | 8% | an entropy written without its minus sign |
| Inverted fraction or reciprocal | 2 | 2% | see below |
| Ambiguous scope | 2 | 2% | cosh−1 meant as 1/cosh |
| Brackets, dimensions | 2 | 2% |
2 of the 118 errors are inverted fractions, which is 1.7% (0.5 to 6.0%). Including ambiguous scope it's 4 of 118, or 3.4%. 974 of the reviewed rows contain a fraction, and 2 of those are inverted.
Inversions
The review found two. The pattern checks found one more in a paper outside the review sample, and one paper has an inversion in its prose.
A 2021 signal processing paper writes a matrix as where its derivation requires . The paper is from before ChatGPT.
A 2025 federated learning paper has this envelope-theorem step:
By the chain rule, the right-hand side should be the other way round. The same section swaps sup and inf
twice, writes max-min for min-max, and has two inequalities pointing the wrong way. Its LaTeX source has
49 U+2010 hyphen characters and 121 inline formulas delimited with \( \) next to the
usual dollar signs. Both counts are the highest among the reviewed papers.
A 2026 materials science paper weights vacancy configurations by , where is the formation energy. The following sentences say the configuration with the lowest formation energy gets nearly all vacancies, which requires .
A 2024 machine unlearning paper defines its pruning score as forget importance divided by retain importance. The appendix calls it "the ratio of the retain importance to the forget importance."
The two ambiguous cases: a 2026 paper on stress rates writes for , which is normally read as arcosh. A 2025 condensed matter paper has , which is only correct when read left to right, while the same paper writes for .
Of these papers, only the 2025 federated learning paper has chat-style characters in its source.
Slashes
The slash check flagged 110 slashes across the 200 papers, and I read all of them. Most are units (mJ/cm²), derivatives (dV/dλ) or sample labels (BN/WSe₂/BN). 14 papers use physics shorthand like ħ²k²/2m or Ω/2π for ħ²k²/(2m) and Ω/(2π). I didn't count those.
Four papers have a slash with unclear scope, for example . One is cs.LG and three are eess.SP. Three of them are from 2026.
The orientation check matched 33 textbook formulas, and all were correct. No paper defined a symbol as both a/b and b/a.
In the review, 4 of 60 papers have an ambiguous or inverted fraction in a display equation: 6.7% (2.6 to 15.9%). The pattern checks on all 200 papers include inline math but only catch specific forms. They find 5 papers (2.5%).
By year
| Year | Reviewed rows | With error | Papers with any error | Chat fingerprint | Unicode em dash | \( math |
|---|---|---|---|---|---|---|
| 2021 | 350 | 4.0% | 7 of 12 | 6 of 39 | 4 | 1 |
| 2023 | 467 | 6.2% | 6 of 12 | 3 of 40 | 2 | 1 |
| 2024 | 302 | 3.6% | 7 of 12 | 8 of 40 | 7 | 6 |
| 2025 | 528 | 7.6% | 6 of 12 | 17 of 40 | 12 | 9 |
| 2026 | 647 | 2.8% | 4 of 12 | 15 of 40 | 11 | 4 |
The error rate varies between 2.8% and 7.6% with no direction over time. Most of the variation comes from a few long papers.
The last three columns count all 40 papers per year. A paper has a chat fingerprint if its source contains
Unicode em dashes, U+2010, U+2011 or U+202F characters, or more than five inline formulas in
\( \). That applies to 6 of 39 papers in 2021, 17 of 40 in 2025 and 15 of 40 in
2026.
Typical LLM words (delve, intricate, showcasing, underscores and similar) have a median of 3.6 per 10,000 words in 2024 and 1.5 in 2026.
The 15 reviewed papers with a chat fingerprint have a pooled error rate of 7.7% and a per-paper mean of 6.1%. The other 45 have 4.2% and 7.2%. The Spearman correlation between LLM word rate and error rate is −0.03.
Limitations
- 60 papers and two inversions, so the intervals are wide.
- The reviewer is an LLM. The rejection pass and the second review don't replace a human referee.
- Inline math was only covered by the pattern checks.
- I used the arXiv source, not the published version.
- Four categories. Theory-heavy fields probably have more equations per paper.
- The data doesn't show why a given equation is wrong.
- I'm not linking the individual papers.
Overall, 4.9% of display equations have a meaning-changing error, 1.7% of those errors are inverted fractions, and neither number changed noticeably between 2021 and 2026.
Sources
- KaTeX contributors. "copy-tex extension." github.com/KaTeX/KaTeX
- Liang, W., et al. "Mapping the Increasing Use of LLMs in Scientific Papers." 2024. arXiv:2404.01268
- Kobak, D., González-Márquez, R., Horvát, E.-Á., Lause, J. "Delving into LLM-assisted writing in biomedical publications through excess vocabulary." Science Advances, 2025. arXiv:2406.07016
- Holzwarth, L., González-Márquez, R., Kobak, D. "Most biomedical publications show signs of LLM-assisted writing." 2026. arXiv:2608.10715
Use your identity on another device
Save an encrypted identity file and restore it in another browser. Keep the file and its passphrase private.