reza

arXiv stats

I had fun exploring trends in arXiv.org papers. Here are some of the figures; if you have comments, contact me.

A few things to look for

  • Who’s on top. The first chart shows CS climbing from fourth place in 2010 to 45% of all submissions in 2025, three times the next field (math).
  • ‘Growth’, but not everywhere. In the counts chart, arXiv as a whole doubles about every seven years and CS about every three. Meanwhile hep-th and hep-ph have barely grown in a decade.
  • Rise and fall. Add your favorite field and follow its share. hep-th was 31% of arXiv in 1993 and is 1.4% today. You can also find your favorite topic, like “dark matter”, in the titles chart and see how common it has been in paper titles over the years (I hope yours made the list :)
  • Changing of the guard. The three biggest categories in 2015 were hep-ph, quant-ph and hep-th. In 2025 they’re cs.CV, cs.LG and cs.CL. quant-ph is the one to watch: its papers nearly doubled between 2020 and 2025, from 6,310 to 11,882 a year, faster than arXiv as a whole.
  • Who lists with whom. The cross-listing chart shows which fields overlap most. In 2025 the biggest pairs were CS with eess, math and stat.
  • The race to be first. The last chart shows when people hit submit. arXiv lists new papers roughly in the order they arrive, so authors pile in right as the daily cutoff passes, to land at the top of the next listing. When arXiv moved the cutoff from 16:00 to 14:00 New York time in 2017, the rush moved with it :)
  • Spot the pandemic. Curious when it hit arXiv? Press Ctrl+F, search “covid”, and jump to its row in the titles chart to see when it became a popular topic.
  • BTW, what’s up with hep-lat: its submissions spike once a year, almost like clockwork, around November and December. Lattice folks, what’s the story?

I hope you have fun with these too.

Hover a line to see its field, click to highlight it, and drag to zoom in (double-click to zoom out). Papers are counted by the month they were first submitted and their primary category.

Each field’s monthly share of arXiv papers

Each field’s monthly count of arXiv papers

Most common topics in paper titles

Topics that were among the 50 most common in titles in at least one year, across all of arXiv or within one area (physics combines every physics archive). A topic is a phrase such as “black holes” or “time series”, several spellings counted together (“LLMs” covers “LLM” and “large language model”), or a single word that names a subject, such as “graphene”; a title counts once per topic. Each row is shaded against the topic’s own peak year, so rising and fading topics stand out whatever their size, and rows are sorted by peak year.

The ten largest fields each year, by rank

Ranked by papers in each calendar year. Follows the Categories / Grouped switch; click a line to highlight it.

How evenly papers spread across arXiv’s fields

This chart measures how evenly each month’s papers are spread across arXiv’s fields, using the Shannon entropy of the field shares, turned into an effective number of fields:

H = −∑i pi ln pi and Neff = eH

Here pi is the fraction of that month’s papers in field i. Entropy is highest when papers are spread evenly and lowest when they sit in one field; eH turns it into a count you can picture. Ten fields with 10% each give Neff = 10. Ten fields where one holds 91% and the other nine 1% each give about 1.6, because in practice nearly everything is in one field. Lower means more concentrated.

The all arXiv line uses every field. The second line repeats the calculation after leaving out the fields you switch off below, with the shares recomputed among the fields that remain. It starts with cs left out, which shows that the recent drop comes from cs. Fields are grouped (cs, math, cond-mat, …) because arXiv added and split categories over the years, which would look like papers spreading out.

Fields in the second line. Click one to leave it out or put it back.

Which fields cross-list with which

Each row is a field, by its papers’ primary category. A cell shows the share of that field’s papers in the year that are also listed in the column’s field, counting a paper once however many of that field’s categories it lists. The bars show the share listed in any other field, next to the average number of categories per paper. Some categories are one subject under two names, such as math-ph and math.MP or cs.SY and eess.SY, so their papers always appear in both. Fields with fewer than 100 papers that year are left out.

Physics papers also listed under machine learning

Each month’s share of physics papers that are cross-listed to cs.LG or stat.ML. A paper counts as physics when its primary category is in a physics archive (astro-ph, cond-mat, gr-qc, hep-*, math-ph, nlin, nucl-*, physics, quant-ph); papers filed mainly under machine learning that cross-list to physics are not included. The all physics line combines every physics archive. Follows the scale and smoothing switches at the top; click a line to highlight it.

How fast each field is growing

How far each month sits above or below its year

Each cell compares a month’s papers with the average month of that year. Empty cells had too few papers.

When papers are submitted

Share of the period’s papers submitted in each hour of the week, in New York time, by when version 1 was submitted. The periods follow arXiv’s changes to its daily deadline, marked by the dashed line: 19:00 New York time from July 1999, 16:00 from October 2001 and 14:00 since 2017.