Guides

How to follow AI research news: where papers actually get published

If you follow AI research news as it happens, you’re mostly reading arXiv. A preprint goes up, the field argues about it within days, and peer review arrives months later at a conference, if it arrives at all. That order matters, because each stage certifies something different. An arXiv posting means a volunteer moderator thought the paper was in scope. A NeurIPS or ICLR acceptance means several experts read it and fought about it in writing, in public. A journal publication means the slowest and most thorough version of that, usually long after the result stopped being news. A lab blog post means a company decided what to tell you.

So the useful skill isn’t finding papers. It’s knowing which checkpoint a given document has passed, so that you can read it accordingly. What follows is the pipeline in order, with the numbers each venue publishes about itself, and the specific things worth checking before you repeat a figure from any of them.

Journals carry the volume, conferences carry the frontier

Start with where the output actually sits, because the answer isn’t what the news cycle suggests. The 2026 Stanford AI Index devotes a chapter to research and development, and its publication section draws on OpenAlex, an open catalogue of scholarly metadata. By that count, AI publications more than doubled between 2013 and 2024, from roughly 102,000 to about 258,000, and AI now accounts for 40.9% of all computer science publications in the database.

Split by venue type, the shape surprises people who only read AI Twitter, though the reason is mundane. Journals take the largest slice, and preprint repositories now sit level with conferences.

Venue typeAI publications in CS, 2024Share
Journal124,00047%
Repository (mostly arXiv)62,52023.7%
Conference62,03023.5%
Book13,0404.9%
Dissertation and other2,3000.9%
Source: AI Index Report 2026, Chapter 1, Figure 1.6.2. The report states the journal and conference shares; the other three shares are our calculation from the same chart.

Two cautions come with that table, and the report raises both. Venue assignment lags, because papers often appear in a repository first and get labelled later, so the newest year understates conferences and journals. And the conference share has been falling for a decade, from 36.6% in 2013 to 23.5%, which reflects growth elsewhere rather than conferences shrinking. The frontier work in language models, agents and reinforcement learning still concentrates in a handful of conferences, which is why a 23.5% slice carries more weight than its size suggests.

arXiv is a timestamp, not a verdict

arXiv is where you’ll read almost everything first, but it has never claimed to be peer reviewed. Its own submission statistics page recorded 3,196,592 submissions in total as of 5 October 2026, across 3,199,023 available articles. The monthly figure is the one that changed character recently, and it’s the one that explains the rest of this section. arXiv received 9,869 submissions in September 2016, 20,569 in September 2024, and 40,363 in September 2026, a record. Submissions doubled in two years, and cs.AI alone rose more than sixfold over that period while other categories roughly doubled.

That volume broke the moderation model, so on 1 October 2026 arXiv capped every submitter at two submissions per calendar month and three active submissions at once, across all categories. Rejected submissions count against the total, because it’s submissions that consume moderator time. September’s 40,363 papers generated almost 9,000 support tickets. arXiv’s moderators report more thin papers of narrow scope, more “salami” papers where one piece of work is sliced into several, and a marked increase in dense, AI-written papers.

However, a relatively small proportion of authors are submitting a large number of low-quality papers and consuming a disproportionate fraction of the moderators’ time. This is unfair to authors who continue to submit quality papers.

Thomas Dietterich, Chair of the arXiv Editorial Advisory Council, via the arXiv blog

The same pressure had already closed one category of arXiv content a year earlier. From 31 October 2025, review articles and position papers in arXiv’s computer science category have needed documented acceptance at a journal or conference before they’re considered at all. arXiv was receiving hundreds of review articles a month, and said the majority were “little more than annotated bibliographies, with no substantial discussion of open research issues”. Technically that wasn’t a policy change, because survey and position papers were never listed as accepted content types. They had just been waved through at moderator discretion, because while there were few of them the cost was small.

Which gives you a precise reading of an arXiv link. A v1 preprint has cleared a scope and plausibility check by a volunteer, and nothing else. That means the author’s institution, the clarity of the method section and the presence of code are doing all the credibility work. If the paper is a survey in a cs category and it isn’t marked as accepted somewhere, it now sits outside what arXiv wants at all.

The conferences do the gatekeeping, and the reviews are public

Machine learning kept its quality control in conferences rather than journals, so the gatekeepers are NeurIPS, ICLR, ICML and the ACL venues for language work. ICLR publishes the most detail about its own process, and its 2026 retrospective is the clearest picture anyone gets of what review at this scale looks like.

ICLR 2026 stageCount
Valid, format-compliant submissions19,525
Desk rejected for procedural or content violations779
Withdrawn before a decision5,042
Received an accept or reject decision13,763
Accepted5,355
Rejected8,408
Reviews written76,139
Reviewers18,054
Source: ICLR 2026 Program Chairs. The chairs state an acceptance rate of 27.4%, which matches 5,355 against all 19,525 valid submissions, not against the 13,763 that reached a decision.

Read that denominator carefully, because it’s where conference acceptance rates get quoted wrong. The chairs report 27.4%, and that figure is 5,355 divided by all 19,525 valid submissions, by our arithmetic. Measured against the 13,763 papers that actually reached a decision, the rate is 38.9%. The gap is the 5,042 authors who withdrew rather than take a rejection, plus the 779 desk rejections. Both numbers are defensible, but mixing them up is how a comparison between two conferences turns into nonsense. The stage counts in the retrospective don’t reconcile exactly either, since 19,525 minus the desk rejections and withdrawals leaves 13,704 rather than 13,763.

The part that helps you most as a reader is ICLR’s openness. Its 2027 program chairs, setting out next year’s submission rules, describe public discussion, manuscript updates during the review phase, and the de-anonymising of even rejected papers at the end of the process as long-standing ICLR practices. That means the reviews sit next to the paper on OpenReview. When a paper’s headline claim seems too clean, the reviewers probably said so, and you can read the authors’ reply.

Conference review is also carrying new load, and the 2026 cycle shows what that costs. ICLR 2026 ran two LLM content detectors across all submitted reviews and flagged entirely machine-generated ones to area chairs. It also ran an automatic reference checker, because many submissions cited documents that don’t exist. Every flagged paper was checked by at least three humans, and confirmed hallucinated references led to desk rejection, which the chairs say partly explains that unusually high 779. On top of that, a malicious user scraped and published the identities of authors, reviewers and area chairs for a large subset of submissions, which prompted collusion attempts and harassment, and forced the chairs to reset all scores to their pre-rebuttal state.

The journal and rolling-review routes, and when they matter

Two venues are worth knowing because they split the thing conferences bundle together. ICLR’s chairs describe review as judging two orthogonal axes: correctness, meaning whether the evidence supports the claims, and significance, meaning whether the finding is new and important. The first is mostly objective and often laborious. The second is entirely subjective.

Transactions on Machine Learning Research deliberately drops the second axis. TMLR says it “emphasizes technical correctness over subjective significance”, runs rolling submissions with shortened review periods, uses double-blind review, and hosts the whole process on OpenReview. So a TMLR paper tells you the method holds up, not that the field found it exciting. That’s a useful signal when you’re checking whether a result is real rather than whether it’s fashionable.

For language work, ACL Rolling Review separates reviewing from acceptance entirely. Authors submit to a cycle, receive reviews and a meta-review, revise in a later cycle while keeping continuity, then commit the reviewed paper to a participating venue for a decision. A paper can therefore carry reviews before it carries a venue name, which is worth remembering when an NLP preprint cites “reviews” without naming a conference.

What a lab blog post leaves out

Then there’s the category that generates most AI research news and sits outside every process above. A company publishes a blog post, a model card or a system card, and the field treats it as a result. These documents are not reviewed by anyone outside the company, and the AI Index is blunt about how much they now withhold, which is the reverse of what the marketing implies.

Industry produced over 90% of notable AI models in 2025, but the most capable models are now the least transparent.

AI Index Report 2026, Chapter 1 highlights

The specifics behind that line: industry accounted for 91.2% of notable models in 2025, against two from academia, and training code, parameter counts, dataset sizes and training duration are no longer disclosed for several of the most resource-intensive systems, including those from OpenAI, Anthropic and Google. Academia still produced 68.1% of AI publications in 2024 and most of the top-cited papers, while industry’s count in the top 100 fell from 17 in 2021 to six in 2024. The people publishing papers and the people building frontier models have come apart.

That gap is why third-party trackers now fill in what releases omit. The AI Index builds its model figures on Epoch AI’s database, which covered more than 3,600 models when we checked it on 5 October 2026 and estimates training compute where labs don’t state it. Epoch is candid that the list is non-exhaustive and manually curated, and the AI Index repeats the warning: it isn’t a census of all models. A chart built on it shows patterns within a curated set, not the whole field. We’ve written separately about why labs release open weights when it suits their business, which is the same decision viewed from the other side.

Reading a paper in ten minutes without getting fooled

A worked example does more here than a checklist, because the failures are specific. “LLMs Get Lost In Multi-Turn Conversation” went up on arXiv on 9 May 2025, and the headline number travelled fast: an average drop of 39% across six generation tasks when a fully specified instruction is split across turns. The ICLR 2026 Outstanding Paper Committee named it one of two Outstanding Papers on 23 April 2026, about eleven months after the preprint by our count. That lag is the normal case, not a delay.

Now read past the 39%. The abstract decomposes the degradation into two parts, “a minor loss in aptitude and a significant increase in unreliability”, drawn from more than 200,000 simulated conversations. Those are different problems with different fixes, so the single percentage hides which one you’re facing. The committee’s own citation adds the limit nobody quoting the number mentions: concerns were discussed about the use of dated models, and the committee judged the conclusions and method still relevant to current systems. That’s a reviewer caveat you can only get by reading the venue’s text rather than the paper’s abstract.

Three habits follow from that. Mapping every headline figure back to the table row it came from catches the cases where an average across tasks and models hides the split. Check which model versions were tested and when, because a 2025 evaluation of 2024 checkpoints is a historical document. And treat benchmark numbers as the weakest part of any paper, for reasons we set out in why AI benchmarks keep lying to you.

There’s also a newer failure mode: the paper may not have been written by its authors. NeurIPS 2026 required position paper submissions to be substantially human-written, then tested every submission with the detector Pangram. At default settings, 273 of 969 submissions, or 28.2%, scored the maximum. Narrowing the detector’s text window to about 100 words cut the share scoring 90% or above from 42.7% to 12.7%, and the chairs went with the narrower setting to avoid over-claiming. The outcome: 178 submissions, 18.4% of the track, were desk rejected, and another 123, or 12.7%, were asked to produce evidence of substantial human engagement or face the same.

Those rates aren’t uniform, which is the part worth carrying away. The same detector flagged 8.2% of 2025 position papers at maximum score against 28.2% in 2026, 2.1% of 2026 submissions to the NeurIPS evaluations and datasets track, and 0% of 159 ACM FAccT papers from 2022, the pre-ChatGPT control. Applied to accepted ICLR 2026 papers, it detected 1%. So argumentative, prose-heavy tracks are where generated text concentrates, and papers that survived review are mostly clean. If you want the same discipline applied to your own citations, that’s the subject of our guide to using AI for research without fake citations.

Where to watch, and what would change this

For day-to-day reading, the practical answer is a small set of feeds rather than a search habit. New arXiv listings in cs.CL, cs.LG, cs.AI and cs.CV carry the raw flow. Hugging Face Daily Papers ranks a subset by community upvotes, which is a popularity signal and not a quality one: the top entry on 5 October 2026 had 93 votes. Conference blogs carry the decisions and the process notes, and the NeurIPS newsletters record things like 102 workshops accepted from 477 proposals for 2026, split across Sydney, Paris and Atlanta. Workshop lists are an underrated early signal, because they show what the field thinks is about to matter before any of it clears review.

The honest limit of all this: nobody outside the labs can verify the frontier claims, and the venues themselves are improvising. ICLR 2027 has rate-limited authors to 20 papers, with a one-paper limit for anyone who has never published at a major AI, ML, CV, robotics or NLP conference. Its chairs argue that review exists to separate signal from noise, and that a field growing faster than its expert pool runs out of qualified eyes. arXiv calls its own two-paper cap a stopgap while it rebuilds moderation tooling. ICLR has also offered submitters access to Google’s Paper Assistant Tool for a one-week window, run outside the review process and invisible to reviewers, after pilots at STOC, ICML and NeurIPS.

Two things would change the reading above. If the submission caps hold, the arXiv firehose narrows, and a preprint starts to carry a weak signal of author curation that it doesn’t carry today. And if automated pre-submission feedback becomes standard, the surface quality of papers rises without the underlying evidence improving, which makes the reviewer discussion, not the manuscript, the thing worth reading. Both are testable within a year, in the venues’ own published statistics.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Leave a Reply

Your email address will not be published. Required fields are marked *