TMLR Editor's Experiment: 3 of 10 Authors Can't Explain Their Own Papers
TMLR editor Nihar Shah interviewed authors of 10 desk-rejected papers, finding three single-author papers whose authors couldn't answer basic questions about problem setup, notation, or where results were supported; two later submitted AI-generated written responses, and one inadvertently described p-hacking, prompting new verification methods like greCAPTCHA.
Background: Submission Flood and Desk Rejection Surge
TMLR (Transactions on Machine Learning Research), launched in 2022 by the JMLR team, operates with volunteer reviewers, action editors, and editors-in-chief. Its review standard emphasizes whether conclusions are supported by evidence, not novelty or SOTA. However, submission volume has tripled in the past year, with single-author submissions increasing 13-fold. Shah notes that desk-rejection rates have risen from about 6% in 2023 to roughly 53% today, meaning over half of submissions never reach external review. Editors have even seen individuals submit five papers in a single day.
The Experiment: Interviewing Authors of Desk-Rejected Papers
During his two-week editorial rotation (August 14–28, 2026), Shah informally selected 10 papers slated for desk rejection. Instead of immediate rejection, he messaged authors via OpenReview, inviting a brief call to better understand the paper before sending it to review. The message read: "An editor-in-chief would like to chat with you before review to better understand this paper; please email available times if convenient."
Response Breakdown
1 author withdrew the paper immediately.
1 author replied they were too busy for a call.
8 authors agreed to meet, but 1 no-showed.
7 authors ultimately met with Shah (undergraduates, master's students, PhD students, faculty, and independent researchers). Most papers were single-author, but not all.
Question Categories
Shah asked two types of questions: (1) basic questions about problem setup, notation, and where in the text the abstract's claimed results are supported; (2) detailed questions about specific technical expressions, theoretical results, and experimental design choices.
Results
Only 1 author answered all questions correctly. However, Shah later found a major error in a key conclusion of that paper; the author acknowledged it. TMLR desk-rejected but allowed resubmission after correction or narrowing the claim.
3 authors could explain the high-level idea but struggled with technical details.
3 authors failed to answer even the basic questions. All three were single-author papers. Shah wrote that two appeared to have virtually no substantive understanding of their own paper's content, and the third could not locate where the abstract's key results were given or supported in the body.
The remaining 9 papers were desk-rejected with no option to resubmit.
Two Revealing Anecdotes
AI-Generated Written Responses
Two authors who couldn't answer basic questions during the call later emailed written responses. Shah ran both through the AI text detector Pangram; both were classified as 100% AI-generated .
Inadvertent Description of p-Hacking
In another meeting, an author tried to explain their analysis method and unintentionally described a full p-hacking workflow — repeatedly adjusting the analysis until the data showed "significance."
Shah's Stance on AI Assistance
Shah explicitly states he is not opposed to AI. He admits using an LLM himself to help read the 8 papers within two weeks, even learning unfamiliar concepts. His concern is not whether authors used AI, but whether they can take responsibility for the content bearing their name.
TMLR's Systemic Countermeasures (Summer 2026)
1. Per-Author Annual Submission Quotas
Unlike fixed per-person limits at conferences, TMLR uses a "harmonic" quota: a paper consumes each author's quota, with the cost decreasing as co-author count increases. Parameters: single-author-only submitters get 2 papers/year; if all submissions are 9-author collaborations, each author gets 9 papers/year. Active reviewers and action editors receive double quotas. Effective July 1, 2026. The design avoids equal splitting to prevent "gift authorship" for quota farming.
2. AI-Assisted Reviewing
Announced July 2026: every submission receives an AI-generated review alongside human review, evaluating only "reliability" — whether conclusions are backed by accurate, clear, convincing evidence. No subjective judgments or acceptance recommendations; final decision remains with the action editor. TMLR selected CSPaper's AI reviewer after evaluation.
3. Clarity as an Acceptance Criterion
On August 28, 2026, TMLR revised its "audience interest" standard to explicitly require papers to clearly communicate findings to readers. The announcement notes that current AI writing tends to be verbose, jargon-heavy, and hard to follow; a paper almost entirely AI-generated with minimal human involvement likely fails this standard with today's AI capabilities.
Broader Landscape: Not Just TMLR
NeurIPS 2026 Position Paper Track
NeurIPS blog (June 2026) reported that of 971 position paper submissions, 273 (28.2%) scored 100% AI-generated by Pangram. The track required papers to be "substantially human-written," allowing AI only for peripheral polishing. AI-generated content growth is pervasive: in the "Datasets and Benchmarks" track, papers with Pangram scores ≥90% increased over 10× from 2025 to 2026. NeurIPS acknowledges detection sensitivity to parameters; using a medium detection window reduced the 90–100% band from 42.7% to 12.7%.
arXiv Policy Change
In May 2026, arXiv CS moderator Thomas G. Dietterich announced a one-year submission ban for all authors of any paper with "clear evidence" of unchecked LLM output (e.g., hallucinated references, leftover chatbot dialogue). Subsequent submissions must first be accepted by a formal peer-review venue. Dietterich emphasized: "Putting your name on a paper means every author is responsible for all its content, however generated."
Biomedical Literature Audit
A Columbia University study published in The Lancet found that the proportion of PubMed Central biomedical papers containing at least one fabricated citation rose from ~0.004% in 2023 to ~0.057% in early 2026.
From Text Detection to Author Verification: greCAPTCHA
Shah acknowledges the interview approach consumed 20–25 hours for only 8 papers, unscalable at current volumes. His team (with CMU's Justin Payan, Bálint Gyevnár, Atoosa Kasirzadeh) proposed greCAPTCHA (named after GRE and CAPTCHA), a method that tests the author rather than the text.
Method
Authors submit their paper and a personal contribution statement. The system auto-generates questions; authors answer under proctored conditions. The output is an evaluation report for journals, hiring, or admissions. The target capability is "capacity to verify" — whether the author can critically evaluate their own contributed content.
Question Types
Identify the paper's actual reported data among multiple versions (error detection).
Explain design choices not justified in the paper.
Explain background concepts the paper assumes known.
Specify concrete conditions under which the method fails.
Pilot Study (31 Researchers)
Each participant answered for their own paper and an unfamiliar paper. Results:
System AUC for distinguishing "own paper" vs. "unfamiliar paper": 0.90 (0.934 without multiple-choice).
Multiple-choice questions alone had near-zero discriminative power (AUC 0.595) because answers are searchable in the PDF.
No statistically significant score difference between in-field and cross-field unfamiliar papers, indicating the test measures more than domain knowledge.
One participant fed the unfamiliar paper to an AI for a full walkthrough beforehand, yet scored 1.3 on it vs. 51 on their own paper.
Current Limitations
Minimum false-negative rate ~20% (1 in 5 genuine authors misclassified).
Participants complained about unclear answer granularity and overly rigid rubrics. One author objected when the rubric rejected their answer for a component they themselves designed.
Co-authored papers: some sections done by co-authors, so authors legitimately cannot answer everything.
Concerns about fairness for non-native English speakers and researchers with disabilities (typing speed, language expression).
Conclusions and Implications
1. Credit Allocation
In today's research ecosystem, academic contribution is largely signaled by authorship. If an author cannot explain or defend their paper, the signal value of a publication as evidence of researcher contribution is severely diminished.
2. Peer Review Cycle at Risk
Journals and conferences routinely ask submitting authors to review others' papers. If authors cannot understand their own work, their ability to serve as competent reviewers is questionable. With AI reviewing and AI writing surging simultaneously, a broken reviewer pipeline undermines the foundation of peer review.
Limitations of the Experiment
Shah notes the sample is tiny (10 papers), informally sampled, and drawn exclusively from papers already deemed desk-reject candidates. It reveals characteristics of low-quality submissions, not the overall TMLR submission distribution. The experiment's purpose was partly to validate the desk-rejection process itself.
Bottom Line
A paper may be completed with AI assistance, but the named author must be able to articulate what is inside it.
Source: Shah, N. B. (2026). "Asking Authors About Their Own Papers." TMLR Blog. https://medium.com/@TmlrOrg/asking-authors-about-their-own-papers-3d2e04e5dee0 NeurIPS Blog (2026). "AI-Generated Papers in the NeurIPS 2026 Position Paper Track." https://blog.neurips.cc/2026/06/02/ai-generated-papers-in-the-neurips-2026-position-paper-track/ Columbia Nursing (2026). "Nearly 3,000 Peer-Reviewed Medical Papers Have Fake Citations." https://www.nursing.columbia.edu/news/nearly-3-000-peer-reviewed-medical-papers-have-fake-citations-columbia-nursing-ai-assisted-audit-finds greCAPTCHA preprint: https://www.cs.cmu.edu/~nihars/preprints/greCAPTCHA.pdf
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
