An AI that can’t write read every AI paper ever written.
TypeSafe’s Jev doesn’t generate text. It answers typed questions with calibrated probabilities, in milliseconds. So we pointed it at every arXiv abstract in cs.AI, cs.CL and cs.LG since 1993 and asked each one the same five questions. Here is what 464,720 papers said.
The ChatGPT-era tells are vanishing from AI papers.
Jev flagged 27.2% of abstracts submitted in Jan 2025 as reading like an LLM wrote or polished them. By Sep 2026 that was 7.4%. Not a cliff, a slide: every month for a year, lower than the last.
It isn’t just the model’s opinion. Plain word counts, computed in code with no AI involved, fall on the same curve. The word novel appeared in 29% of abstracts at its peak and sat between 15% and 25% for fifteen years. It is now in 6%. Underscores went from 4.1% to 0.7%. Delve is gone.
Two explanations fit. Either researchers stopped letting AI write their abstracts, or the newer models stopped sounding like AI and started scrubbing the old tells. The data can’t tell those apart. Neither can you.
One in four AI papers in 2025 read like an AI wrote them.
Zoom out to years and the ChatGPT inflection is unmistakable. 6% of 2022 abstracts read LLM-written. In 2025 it was 24%. Then 2026 falls back to 12%, and that’s with the year only three-quarters over. The field wrote its own papers with the tool it was studying, then either stopped or got better at hiding it.
Code release went from a rounding error to one paper in six.
In 2010, 0% of abstracts said code, models or data were public. By 2024 it was 17%. That’s the reproducibility movement, GitHub and Hugging Face showing up in the text of the science itself. It also means five in six papers still don’t mention releasing anything.
Theory collapsed. Benchmarks exploded.
Jev sorted every abstract into one of six kinds. Theory papers were 25% of the field in 2008 and 6% in 2025. Benchmark and dataset papers went from 3% in 2019 to 12% in 2026, the fastest-growing kind of AI paper. Method papers hold steady around 65%: the field still mostly proposes things.
“State of the art” never got more common. The hype did, then didn’t.
About 11% of abstracts claimed a state-of-the-art result in 2016. In 2025, 15%. For all the leaderboard culture, the share of papers claiming to top one has been flat for a decade. Promotional language is a different story: Jev’s hype score climbed steadily through 2025 and is now falling in lockstep with the LLM tells above.
The abstracts Jev was most sure about.
Most promotional language
- 1deepTarget: End-to-end Learning Framework for microRNA Target Prediction using Deep Recurrent Neural Networks20161.00
- 2We Built a Fake News & Click-bait Filter: What Happened Next Will Blow Your Mind!20181.00
- 3MemGEN: Memory is All You Need20181.00
- 4Revolutionizing Single Cell Analysis: The Power of Large Language Models for Cell Type Annotation20231.00
- 5Alternative Telescopic Displacement: An Efficient Multimodal Alignment Method20231.00
- 6Personalized Resource Allocation in Wireless Networks: An AI-Enabled and Big Data-Driven Multi-Objective Optimization20231.00
- 7A Revolution of Personalized Healthcare: Enabling Human Digital Twin with Mobile AIGC20231.00
- 8Digital twin brain: a bridge between biological intelligence and artificial intelligence20231.00
Most likely LLM-written (2023+)
- 1A Novel Approach to Breast Cancer Histopathological Image Classification Using Cross-Colour Space Feature Fusion and Quantum-Classical Stack Ensemble Method202495%
- 2Edge AI for Internet of Energy: Challenges and Perspectives202394%
- 3Deep Learning and Machine Learning, Advancing Big Data Analytics and Management: Unveiling AI's Potential Through Tools, Techniques, and Applications202494%
- 4State-of-the-art AI-based Learning Approaches for Deepfake Generation and Detection, Analyzing Opportunities, Threading through Pros, Cons, and Future Prospects202594%
- 5Ethical Framework for Harnessing the Power of AI in Healthcare and Beyond202393%
- 6A Study on the Implementation of Generative AI Services Using an Enterprise Data-Based LLM Application Architecture202393%
- 7Deep Residual CNN for Multi-Class Chest Infection Diagnosis202393%
- 8Exploring the intersection of Generative AI and Software Development202393%
Least LLM-sounding recent papers (2025+)
- 1Relative Prime Factorization and Finite-State Presentations under Fixed Finite-Monoid Observation20269%
- 2Minimal Sequent Calculus for Teaching First-Order Logic: Lessons Learned202510%
- 3Generalizing Analogical Inference from Boolean to Continuous Domains202510%
- 4A solution to the Erd\H{o}s Problem #1040202610%
- 5Edge-addition monotonicity of positive p-energy fails for every p >= 1202610%
- 6Language Independent Named Entity Recognition via Orthogonal Transformation of Word Vectors202511%
- 7Universal Coefficients and Mayer-Vietoris Sequence for Groupoid Homology202611%
- 8Run-and-tumble chemotaxis using reinforcement learning202512%
Most confident SOTA claims
- 1PerfectDou: Dominating DouDizhu with Perfect Information Distillation202298%
- 2unMORE: Unsupervised Multi-Object Segmentation via Center-Boundary Reasoning202598%
- 3Solving Formal Math Problems by Decomposition and Iterative Reflection202598%
- 4A Neural Autoregressive Approach to Collaborative Filtering201697%
- 5Relationship-Embedded Representation Learning for Grounding Referring Expressions201997%
- 6Egocentric Object Manipulation Graphs202097%
- 7COVIDX: Computer-aided diagnosis of Covid-19 and its severity prediction with raw digital chest X-ray images202097%
- 8Neural Sentence Ordering Based on Constraint Graphs202197%
Scores are model judgments about abstract text, not about the papers’ quality or truth. Click a title to read it and decide for yourself.
How an AI that can’t write read 464,720 papers.
Jev is TypeSafe’s System One model. It doesn’t produce prose. You give it state and typed questions, and it returns a probability for a yes/no, a choice from a fixed set, or a position on a rubric, each calibrated. Output tokens are free and input is $0.042 per million, which is why this cost $11.20 instead of a few thousand dollars with a chat model.
Source: the Cornell arXiv metadata snapshot on Kaggle (CC0), filtered to any paper listing cs.AI, cs.CL or cs.LG, primary or cross-listed: 464,720 papers, 1993–2026. Sixteen abstracts went into each request as a JSON array, with five questions per abstract referencing it by index. All 464,720 papers are scored. The run happened in two sessions on two accounts' starting credit, 27.6 minutes of wall clock in total.
The five questions, verbatim:
- Noul · claims state-of-the-art or beating all prior methods
- Noul · says code, models or data are publicly released
- Noul · abstract reads as written or polished by an LLM: formulaic phrasing, uniform rhythm, words like delve, underscores, paving the way
- Choice · kind of paper: method, benchmark, survey, theory, application, position
- Score · promotional hype level: plain technical wording / some emphasis words such as significant or novel / strong promotional wording such as revolutionary, paradigm shift, unprecedented
Caveats, because there are real ones. The LLM-written question describes 2023-era tells, so a drop can mean either less AI writing or AI writing that no longer has those tells; the word-count lines are there precisely to show the shift is in the text, not the judge. Yes/no shares use a 0.5 probability threshold. Years with fewer than 100 scored papers are omitted. 2026 covers January through September. Jev judges what an abstract says, not whether it’s true.
Everything is reproducible: filter script, batched runner, analysis and this site are in the repo. Total wall-clock for the scoring run was 27.6 minutes at roughly 281 papers per second, bounded by the API’s rate limit, not the model.