ArXiv, a widely used open-access repository, has begun aggressively cracking down on the careless use of AI large language models in scientific papers. Although papers are uploaded to the platform before being peer-reviewed, arXiv remains one of the primary ways research circulates in fields such as computer science and mathematics. The site itself has also become an important source of data on trends in scientific research.
To address the growing number of heavily AI-generated papers, arXiv has introduced stricter measures, including requiring first-time authors to obtain endorsements from established researchers before posting submissions.
Thomas Dietterich, chair of arXiv’s computer science section, wrote on Thursday that “if a submission contains incontrovertible evidence that the authors did not check the results of LLM generation, this means we can’t trust anything in the paper.”
He added that evidence of AI misuse could include hallucinated references, comments addressed to or generated by the LLM, and other obvious signs of unverified AI-generated content. Dietterich further stated that if such inputs are found in a paper, authors could face “a 1-year ban from arXiv followed by the requirement that subsequent arXiv submissions must first be accepted by a reputable peer-reviewed venue.”
These policy changes underscore arXiv’s insistence that authors take “full responsibility” for their content, “irrespective of how the contents are generated.” This means researchers remain accountable if they directly copy and paste “inappropriate language, plagiarized content, biased content, errors, mistakes, incorrect references, or misleading content” from an LLM into their papers.
Dietterich also told media outlets that this would function as a “one-strike” rule. However, moderators would first need to flag the issue, and section chairs would have to confirm the evidence before any penalty is imposed. Authors will also have the right to appeal decisions.
Recent peer-reviewed research has additionally found that fabricated citations are increasingly appearing in biomedical studies due to the misuse of large language models.








