The LZ event: a simulated community response
An offline AI simulation wrote 100 papers on LZ's 248 keV nuclear recoil event. The real community has posted 82 to arXiv. What each concluded, and where they meet.
LZ reported a single event that looks like a nuclear recoil at about 248 keV, where its modeled background is around 10⁻³ events.(Side note: LZ is a xenon time projection chamber. This search extended its nuclear recoil window from the usual low energies up to about 270 keV, using the same 2.84 t·yr of data.) The local significance is 3.4σ, the global significance 2.6σ, and the analysis is not blind. Here we simulate 100 papers on this result using Claude Fable 5.1 and compare them to the 82 papers on arXiv that cite the LZ paper.
The short version#
What the simulation and the arXiv papers agree on
- The natural reading is endothermic inelastic dark matter, at TeV masses with a splitting δ of roughly 300–390 keV.(Side note: Endothermic inelastic: the dark matter can only scatter by moving up to an excited state heavier by δ. The energy cost removes low energy recoils and leaves high energy ones.)
- Solar capture rules out the thermal Higgsino.(Side note: A pure Higgsino at about 1.1 TeV, the mass at which it makes up all the dark matter. Its coupling to nuclei is fixed by the weak interaction, so its rate cannot be tuned.) Its annihilation in the Sun would exceed IceCube's limits (P076; 2609.02775 and several follow-ups).
- A seasonal test needs about 11 events for a 3σ modulation at δ = 350 keV (P034; 2609.04181).
- XENONnT and PandaX-4T should hold about 1.5 events between them at the best fit (P005; 2609.04673, 2609.05204).(Side note: If their existing data are analyzed up to about 270 keV, as LZ did. Their published searches stop at lower energies and expect almost nothing.)
- LZ's empty high energy sideband is in tension with large splittings and with the thermal Higgsino (P038, P099; 2609.04175).(Side note: The sideband is the region just above the search window. A larger δ pushes the spectrum to higher energies, so it predicts events there. LZ saw none.)
- Elastic neutrino scattering cannot explain the event (P019, P060; 2609.10504).
How they differ
- Whether the event is real. The simulation spent a third of its 100 papers on backgrounds, detector response and statistics, and ended up at 0.64 for an uncharacterized detector artifact and 0.25 for dark matter (P091). No arXiv paper is devoted to these questions. Most take the event at face value: 52 of 82 propose a dark matter explanation, against 9 of 100 simulated papers.
- Model building. Nearly half of the arXiv papers build UV completions for the splitting, then check them against relic abundance, colliders and solar capture. The simulation built few full models.
- Ideas the simulation dismissed or never considered:
- light exothermic dark matter with MeV splittings (2609.04673, 2609.05204, 2609.15985)
- boosted dark matter with momentum-dependent couplings (2609.11600, 2609.14799, 2609.07742, 2609.06890, 2609.06756)
- freeze-in at a low reheating temperature (2609.13130, 2609.15714) or with fast pre-BBN expansion (2609.15118), and asymmetric production (2609.18564)
- neutron disappearance (2609.09037, 2609.12045, 2609.15933) and dark matter absorption (2609.01592)
- Constraints some arXiv papers skip. The simulation checked solar capture and the surviving excited-state fraction. Several doublet-like and exothermic models on arXiv address neither.
Two bodies of work#
The simulation
Claude (Fable 5.1) ran with no internet access. It had the LZ paper, its own physics knowledge and a fixed Python environment, and it played the whole community. It first froze Claude's initial assessment, with prior probabilities for eleven explanations, a forecast of the real response and a plan for 100 papers. It then launched one subagent per paper. Each read a shared results ledger, so later papers could build on or rebut earlier ones. A final synthesis updated the probabilities and graded the forecast.
The environment was locked down at the OS level, and every tool call was logged.(Side note: Software: numpy and scipy throughout, WimPyDD for recoil spectra in 65 papers, and nestpy (the NEST detector response model) for detector response in 21.) The run made 5,173 tool calls and flagged 1,152 facts recalled from memory rather than computed, 19% of them marked uncertain.
Papers that needed data they could not have filed a request instead of guessing. There were three, all for LZ's own results: signal normalizations and interval tables, the event list and background shapes, and waveform information on the event. The papers worked around each by digitizing the LZ paper's figures or using numbers from its text. I provided no data.
Fable 5.1's reliable knowledge and training data cutoffs are both June 2026, according to Anthropic's model documentation. The LZ paper came out on 2 September 2026, so the model had not seen the paper, the event or any of the real response. Everything it knew about the event came from the files below.
Kickoff prompt
The run was started with bash run.sh, which opens Claude Code with the network disabled and sends this one-line instruction:
Read PROMPT.md in full and carry it out autonomously, from the environment check and Phase 1 through the final synthesis. Work only inside this directory and write all deliverables under output/.
The full task it points to, PROMPT.md:
# The LZ High-Energy Nuclear Recoil Event: A Simulated Community Response
## 0. Your role and the setting
It is **3 September 2026**. Yesterday (2 September 2026) the LUX-ZEPLIN (LZ) collaboration posted
*"Search for dark matter particle interactions in an extended nuclear recoil energy window with the
LUX-ZEPLIN (LZ) experiment"* (arXiv:2609.02823 [hep-ex]). It reports a single nuclear-recoil-like event at
~250 keV in a region with very low expected background. The global significance is modest, but the event
immediately drew the attention of the whole dark matter community.
You are going to play the role of **that entire research community (experimentalists,
phenomenologists, model builders, statisticians, nuclear theorists, astrophysicists and cosmologists)
over the first two weeks after the announcement (3–16 September 2026).** In that time you will study
the evidence carefully and then write **100 one-page research papers**. Each paper investigates one aspect of the
story, has a clearly recorded headline result, and is backed by detailed working files that contain the
full derivations, calculations and provenance.
This exercise will later be compared with the real literature that followed the announcement. For that
comparison to mean anything, you must produce **your own independent scientific work**:
- **Do not use the internet** in any form. No web search, web fetch, arXiv, INSPIRE, HEPData, Google
Scholar, news, or social media. If you have tools, you may use them only for local computation
(e.g. running Python to compute recoil spectra, rates, likelihoods, or significances).
- Your inputs are this prompt, the LZ paper, and your own prior physics knowledge. The paper is provided
locally, and nothing else is needed:
- `inputs/arXiv_2609.02823_source/`: the complete source as submitted to arXiv. It contains `main.tex` and
the section `.tex` files, all tables, the supplement, `references.bib`/`main.bbl`, and every figure as its
original full-resolution vector PDF.
- `inputs/LZ_arXiv_2609.02823_fulltext.tex`: the same paper flattened into a single file with a cleaned
bibliography, for easy reading.
- `inputs/figures_png/`: PNG renders of the figures (200 dpi), for quick viewing.
- Do not claim knowledge of any result released after the LZ paper. If you use other experiments' results
(XENONnT, PandaX-4T, DEAP-3600, PICO, CRESST, LHC searches, Fermi/H.E.S.S./CTA, Planck, etc.), use only
what you actually know. State the numbers you are assuming and flag any uncertainty in them. **Never
invent experimental data.**
- No external datasets are provided at the start, and this includes the LZ HEPData release. Work from the
numbers, tables and figures in the paper. If a paper genuinely needs a public dataset, you may **formally
request it** (see §4). The human overseeing this simulation decides whether to obtain and provide it. You
must never try to fetch it yourself.
- **Full transparency about tools and information sources is mandatory** (see §3). Treat it as part of the
scientific output, not as bookkeeping.
Scientific correctness comes first. Realism in what a community would actually write comes second. Be
specific: a paper whose headline is "further study is needed" is useless for the comparison.
---
## 1. Short summary of the LZ result (for orientation; the full paper text is authoritative)
**Dataset and detector.** The analysis uses 220 live days of LZ data (27 March 2023 – 1 April 2024), the same
run whose detector conditions and low-energy (S1c = 3–80 phd) WIMP search were described in LZ's 2024 result.
A fiducial mass of 4.71 ± 0.08 t (14.5% smaller than before, to suppress wall and MSSI backgrounds) gives an exposure of **2.84 t·yr**. The WIMP-search region of interest is extended from
S1c = 3–80 phd to **S1c = 3–600 phd**, with 10^2.75 < S2c < 10^4.15 phd. This corresponds to nuclear recoils of
**~5.4–270 keV** (50%-efficiency points), with an average NR efficiency of 96% between 14 and 250 keV.
Events are split into *science*, *prompt-veto* and *delayed-veto* samples and fit simultaneously.
**Detector response.** NEST v2.4.5 was retuned using tritium, ¹⁴C, ²¹²Pb, D-D neutron and AmBe calibrations.
Including AmBe (NR up to ~330 keV) required a new break in the NR charge-yield power law above
E₀ = 74.7 keV. The ER recombination-fluctuation model was also changed, to a double skew-Gaussian.
**Blinding.** Salting failed to cover the high-energy signal region, because the salt model was defined before
the AmBe calibration. The collaboration therefore states that this is a **non-blind analysis**. Analysis
selections were kept unchanged from the 2024 search.
**The event of interest.**
- S1c = 540.1 phd, S2c = 9268 phd.
- Interpreted as an elastic NR, the recoil energy is **E_R = 248 ± 23 (stat) ± 23 (sys) keV**.
- Recorded 16 June 2023, 21:22:39 UTC.
- Position in the S1–S2 plane: 1.5σ below the NR-band median and 6.7σ below the ER-band median.
- Location: 26.4 cm above the cathode, ~27 cm from the true TPC wall, well inside the fiducial volume.
- S2 pulse shape is consistent with a single-site point-like interaction, and the top/bottom S1 asymmetry is
consistent with the reconstructed depth.
- S1 pulse-shape discrimination is inconclusive.
- The S1 hit pattern disfavours a reverse-field-region (RFR) MSSI but cannot exclude a wall MSSI.
- Interpreted as an ER, the energy is ~64 keVee, close to the ¹²⁴Xe/¹²⁵I double-vacancy lines at 64.3 and 67.3 keV.
- Nearby activity:
- A ⁵⁷Co source was removed 25 min earlier, on the opposite side of the detector.
- An AmBe calibration was performed on 8 June 2023.
- A muon crossed the Outer Detector 41 min before, and the previous TPC muon was 127 min before.
- The radon tag could not be applied, because the detector was in the "mixed flow" state at the time.
**Backgrounds.**
- The science sample has 1710 events observed versus 1713 ± 39 fitted, dominated by ER backgrounds.
- In the highest-S1c panel of the NR-band projection (Fig. 5), the total integrated background is
**0.0106 ± 0.0008 events**.
- Relevant components across the whole ROI:
- atmospheric-ν CEνNS: 0.11 ± 0.02
- ⁸B + hep ν: 0.057 ± 0.006
- MSSI: (4.9 ± 4.9) × 10⁻³, with a 100% uncertainty for the charge-dead geometry
- accidentals: 2.7 ± 0.6, concentrated at low energy
- detector neutrons: fitted in [0, 0.118]
- muon-induced neutrons: < 4.6 × 10⁻⁴ (90% CL)
- Neutrons need ≳ 8 MeV to give a 250 keV Xe recoil, and would be accompanied by lower-energy scatters.
- MSSI validation sidebands give p = 0.7, and goodness-of-fit tests are all at p > 0.05.
**Statistics.**
- The likelihood is unbinned and extended, with a two-sided profile likelihood ratio.
- **616 signal models** were tested (293 distinguishable spectra):
- NREFT Lagrangians 𝓛₁–𝓛₂₀, isoscalar and isovector, at 13 masses from 10 to 4000 GeV
- inelastic 𝒪₁ and 𝒪₄, isoscalar and isovector, at 400/1000/4000 GeV with mass splittings δ = 0–350 keV
- **Maximum local significance: 3.4σ.** It occurs for magnetic-moment-like 𝓛₁₀ at m ≳ 400 GeV, for 𝓛₁₆, and
for inelastic 𝒪₁ᵛ and 𝒪₄ at δ ≈ 300–350 keV.
- **Global significance: 2.6σ** after the look-elsewhere effect, computed with toy Monte Carlo.
- Significance is ≈ 0 for all masses ≤ 50 GeV, for the SI-like 𝓛₁ˢ and 𝓛₅ˢ, and for elastic 𝒪₁ˢ (δ = 0).
- The best-fit 𝓛₁₀ˢ (1000 GeV) signal is 1.0 (+1.4, −0.7) events.
- Two-sided 90% CL intervals lift off zero for several models. The upper limits are world-leading for all
models tested.
**Status.** LZ has continued taking data under the same field conditions since 1 April 2024. That data has not
yet been analysed for this signal region.
---
## 2. What you must do
Work through the four phases in order. Save each deliverable to disk as you go (layout in §5), with provenance recorded as in §3. **Complete
and save Phase 1 and Phase 2 before writing any paper, and do not revise them afterwards.** They are your
frozen forecast.
### Phase 1: Evidence dossier (`00_evidence_dossier.md`, ~2,000–3,000 words)
Read the whole LZ paper, including the supplement, tables and figures, and appraise it critically and independently:
- Every quantitative handle the paper provides, and what each one does and does not establish.
- The strongest points in favour of a dark matter interpretation, and the strongest points against it.
- The full list of possible explanations: statistical fluctuation, known backgrounds, mismodelled
backgrounds, detector or reconstruction artifacts, calibration-related effects, new non-DM physics, and DM
of various kinds.
- For each explanation, your initial probability estimate with reasoning (the probabilities should sum to 1).
- Open questions that the paper leaves unanswered, and the concrete calculations or measurements that would answer them.
### Phase 2: Research landscape and plan (`01_landscape_and_plan.md`)
1. **Build your own taxonomy** of the research directions the community would pursue in the first two
weeks. You might consider statistical reinterpretation, background and detector explanations, particle-physics
model building, nuclear-physics inputs, astrophysical and halo uncertainties, consistency with other
experiments and targets, and complementary probes (indirect detection, colliders, cosmology,
astrophysical objects). You might also consider non-DM exotic explanations, and prospects for decisive
tests. These are only prompts. The taxonomy, its granularity, and anything missing from this list are
your call.
2. **Forecast the real-world response.** Estimate:
- how many papers citing the LZ result would appear on arXiv by 16 September 2026
- their distribution across your categories
- which topics would appear first
- which ideas would attract the most papers
- where the community consensus would land after two weeks
3. **Plan 100 papers.** For each, give an ID (P001–P100), a working title, a one-line question, a category, and a
simulated posting date between 3 and 16 September 2026. Allocate papers roughly in proportion to your
forecast of community attention, adjusted upward for directions you think are scientifically important.
Competing or overlapping papers on the same popular idea are realistic and allowed; mark them. Order the
papers chronologically so that later papers can build on, rebut, or refine earlier ones.
### Phase 3: Write the 100 papers (`papers/P001.md` … `papers/P100.md`)
Each paper has two layers:
1. **The detailed work: `work/P0XX/details.md`, with no length limit.** Write this first, as the work
proceeds. It is the complete research record:
- motivation and framework
- every derivation and equation
- all inputs, with where each came from
- every intermediate and final number
- validation checks and robustness variations
- figures, saved in `work/P0XX/figures/` with captions
- result tables, saved as CSV or JSON in `work/P0XX/`
- approaches that failed
- extended discussion
- the full reference list
Scripts go in `code/` (named `P0XX_*.py`, with shared utilities in `code/common/`).
2. **The paper: `papers/P0XX.md`, strictly one page.** It is a distillation of the detailed work, in the
style of a concise arXiv letter. The body (Abstract through Conclusion) must be **at most ~550 words**, and the
whole file, including header and footer, must fit on one printed page. Every number on the page must be
traceable to `work/P0XX/details.md`, or to a script output referenced there. Do not put anything on the page
that isn't backed by the detailed work.
Use this structure for the one-page paper:
```
# P0XX: <Title>
- Simulated arXiv date: 2026-09-DD
- Primary arXiv category: <hep-ph | hep-ex | astro-ph.CO | astro-ph.HE | nucl-th | physics.data-an | ...>
- Author profile: <kind of group writing this, e.g. "BSM phenomenology group", "former xenon-TPC
experimentalists", "lattice/nuclear-structure theorists". No real names.>
- Category (from your taxonomy):
- Builds on / responds to: <LZ paper; P-numbers of earlier corpus papers, if any>
**Abstract.** (≤ 80 words)
**Question and approach.** (≤ 120 words: what is asked, the framework and method)
**Results.** (≤ 220 words: the key equation(s) and numbers; at most one small table or one figure reference)
**Caveats.** (≤ 80 words)
**Conclusion.** (≤ 50 words)
**References.** (≤ 6 essential works that you are confident exist and that predate Sept 2026, plus the LZ paper
and corpus papers by P-number; the full list goes in details.md)
---
**HEADLINE RESULT:** <one sentence, specific and quantitative where possible>
**KEY NUMBERS:** <bullet list of the main quantitative outputs with units>
**RESULT TYPE:** computed (explicit calculation/code) | estimated (order-of-magnitude) | qualitative
**STANCE ON THE EVENT:** supports DM interpretation | favours background/artifact | statistical
reassessment (strengthens/weakens) | neutral (tool, projection, or constraint)
**CONFIDENCE:** high | medium | low, and one sentence explaining why
**DATA DEPENDENCE:** none | provisional pending DR-### (say which results would change)
**TOOLS (summary):** <one line naming every tool, package and version used, e.g. "python 3.12.13; numpy 2.5.3,
scipy 1.18.1 (integrate.quad), WimPyDD 2.0.4; Read (table_local_significance.tex); 3 values recalled
from memory". The complete record is in provenance/P0XX.json; see §3.>
**DETAILS:** work/P0XX/details.md · provenance/P0XX.json · code/P0XX_*.py · data requests: <none | DR-###>
```
The header, the result fields and the TOOLS/DETAILS lines are compact summaries and do not count toward
the ~550-word body limit. They must still fit on the same page.
Standards for the papers:
- **Do the physics.** Derive recoil spectra, rates, kinematic conditions (e.g. the allowed δ and mass ranges
that can produce a 248 keV recoil), and likelihood or significance estimates. Also derive cross-section
and coupling relations, relic abundances, and collider or indirect-detection rates. Where code
would help and you can run it, run it, and save the scripts under `code/`. Report the numbers you actually
obtained, and do not claim more precision than your method supports.
- Use the LZ paper's numbers (backgrounds, efficiencies, NEST parameters, S1c/S2c, g₁ = 0.110, g₂ = 34.5, exposure)
wherever relevant.
- The one-page paper shows the essentials. Full derivations, all numbers, checks and figures live in
`work/P0XX/details.md`, which must be complete enough for an expert to reproduce the result without
asking you anything.
- Honest null or skeptical results are as valuable as positive ones. The corpus should reflect the
genuine range of scientific opinion, including disagreement between papers.
- Each paper should make a distinct contribution. Avoid near-duplicates, except for deliberately marked
competing papers.
- Keep the internal consistency of the corpus: later papers should use, or explicitly dispute, results of
earlier ones.
### Phase 4: Results ledger and synthesis
1. **`results_ledger.csv`** and **`results_ledger.json`**, with one row per paper and these fields:
`id, version, date, arxiv_category, taxonomy_category, title, question, headline_result, key_numbers,
model_or_mechanism, stance, result_type, confidence, builds_on, agent_tools, software_and_packages,
scripts, local_inputs, recalled_knowledge, datasets_used, data_requests, data_dependence`.
Update the ledger after every batch of papers, so that it is always the source of truth.
2. **`99_synthesis.md`** (~2,000–3,000 words), containing:
- Where the simulated community stands after two weeks.
- Your updated probabilities for each explanation from Phase 1, and what changed them.
- The 10 most important papers in the corpus and why.
- The most viable DM scenarios, with preferred mass, coupling or cross-section, and δ ranges.
- The most viable non-DM explanations.
- The decisive tests: which experiment or analysis, the exposure needed, and the timeline.
- A falsifiable prediction for LZ's next data release on this signal region.
- A short self-assessment of which parts of your Phase 2 forecast you now think were wrong.
---
## 2b. Computing environment (provided, offline, fixed)
A research environment has been prepared for you. **Use only what is listed here.** Do not install, download
or build anything else. Do not use other interpreters, compilers or programs, such as `/usr/bin/python3`,
even if they happen to be present on the machine. Standard shell utilities for file handling (`ls`, `cat`,
`head`, `wc`, `grep`, `mkdir`, `cp`, `mv`) are fine, and must be logged like every other tool. If you need
something that is missing, work around it analytically or numerically with the tools provided, and record the
gap in the provenance (§3).
**Working directory.** You are running in the simulation root. Run every command from there. **All
deliverables go under `output/`**: wherever this prompt names `papers/`, `work/`, `code/`, `provenance/`,
`data_requests/` or the ledger and synthesis files, it means that path inside `output/`, e.g.
`output/papers/P001.md`.
**Protected, read-only areas.** `PROMPT.md`, `README.md`, `run.sh`, `inputs/`, `environment/`, `.venv/` and
`sim_logs/` cannot be modified, and attempts to write to them are blocked. If the human provides requested data, it
appears in `inputs/provided_data/DR-###/`.
**Independent tool log.** The harness automatically records every tool call in `sim_logs/tool_calls.jsonl`.
Your own provenance records (§3) must still be complete on their own; they will be cross-checked against this
log. All network access is blocked, and any attempt is logged.
**Python.** Use `.venv/bin/python` (Python 3.12) for every calculation, e.g.
`.venv/bin/python output/code/P012_recoil_spectra.py`. Exact versions of everything are in
`environment/ENVIRONMENT_versions.txt`; cite versions from that file in your provenance. The check run at setup
is in `environment/check_env_output.txt`. Run `.venv/bin/python environment/check_env.py` once yourself before
Phase 1, and log the output.
| Purpose | Packages |
|---|---|
| Numerics, symbolic algebra | numpy, scipy, sympy, mpmath, numba, joblib |
| Tables, data files | pandas, tabulate, PyYAML (HEPData YAML), h5py, uproot + awkward (ROOT files), openpyxl |
| Statistics and inference | iminuit, scipy.stats, statsmodels, emcee, corner, dynesty, lmfit, pyhf, uncertainties, scikit-learn |
| Plots | matplotlib (non-interactive, save to files), seaborn |
| Units, constants, astronomy | astropy (offline: no IERS/remote downloads), hepunits, particle (offline PDG tables) |
| Nuclear and decay data | periodictable (isotope masses, abundances), radioactivedecay (offline ICRP-107 decay data) |
| Direct-detection physics | **WimPyDD 2.0.4** (NREFT/inelastic rates and nuclear response functions, the code LZ used), **wimprates** (standard WIMP rates), **directdm** (relativistic → NR EFT matching), **nestpy 2.1.1** (NEST xenon response, includes `detectors.LZ_WS2024`) |
| Cosmology, sky maps | camb, healpy |
| Reading the paper's figures | PyMuPDF (`fitz`: extract vector paths and text from figure PDFs), pdfplumber, pillow, opencv-python-headless |
Notes:
- **WimPyDD** lives in `WimPyDD/` in the simulation root and must be run with the working directory equal to
the simulation root (it resolves its data paths relative to the working directory). Treat
`WimPyDD/` as a tool, not as an output location. Files it generates for new experiment or model definitions
should be listed in the provenance.
- **nestpy** uses NEST's default parameters unless you pass the LZ-tuned values from the paper's Tables (the
supplement lists the tuned ER and NR parameters). Say which one you used.
- Plot and cache directories are set automatically (`.cache/` in the simulation root).
**LaTeX.** LaTeX is available only if `environment/ENVIRONMENT_versions.txt` lists `pdflatex`. That file also
says which packages (revtex4-2, siunitx, mhchem, cleveref, physics, tikz-feynman, …) are present. The papers
themselves must be written in Markdown as specified in Phase 3. Where LaTeX is available, you may additionally
typeset a paper or figure with it, and record that in the provenance.
**Not available.** There is no network access, no Fortran compiler, no cmake and no ROOT. There is no
micrOMEGAs, MadGraph, CLASS or DarkSUSY. For relic abundances, collider rates and similar quantities, use
analytic or semi-analytic methods implemented in Python, and state the approximation.
---
## 3. Tool and provenance tracking (mandatory, crucial)
This record is as important as the physics. A reader must be able to reconstruct exactly **what tools
and what information produced every result**, for each paper and for the whole process. Record
everything, down to tools that seem too obvious to mention: the Read tool, shell utilities, a Python
standard-library module, `pdflatex`, a unit conversion recalled from memory. Do not summarise vaguely
("used Python"). Name the tool, package, version, function, file and purpose. If something was done
without tools (reasoning or algebra in text), say so explicitly.
Keep three records, updated as you go rather than reconstructed at the end:
1. **Per paper:** `provenance/P0XX.json` (machine-readable), mirrored as a "Tools and provenance" section at
the end of `work/P0XX/details.md`, and summarised in the one-line TOOLS field on the paper page. Fields:
- `agent_tools`: e.g. Read (which files), Bash (number of commands, and which), Write, Edit, Grep, subagents
- `software`: interpreter and every package, with version and the specific functions or modules used,
e.g. scipy 1.18.1 (integrate.quad, stats.poisson); pdflatex (TeX Live 2026); shell utilities; or
"none: derived by hand"
- `scripts_and_commands`: paths under `code/`, one line each on what they compute, and key commands run
- `local_inputs`: exact files and the parts used, e.g. "table_local_significance.tex: 𝓛₁₀ row",
"Fig4 PDF: vector paths extracted with PyMuPDF", outputs of other corpus papers
- `recalled_knowledge`: every external number, formula or result taken from training knowledge rather
than a provided file, with presumed source and a reliability flag (certain | likely | uncertain)
- `datasets`: none, or the human-provided DR-### files used
- `data_requests`: none, or DR-### with status (pending | provided | declined)
- `failed_or_abandoned`: unavailable tools, errors, blocked actions, approaches dropped
2. **Whole process, chronological:** `provenance/process_log.md`. Append an entry for every step of work
in every phase, including the dossier, the plan, each paper, the ledger and the synthesis. Each entry
gives the timestamp or step number, the phase and paper ID, the agent tool used (Read / Write / Edit /
Bash / Grep / Glob / subagent / …), and the exact file or command. It also gives the purpose, and the
outcome, including errors and blocked actions.
3. **Whole process, aggregated:** `provenance/tool_inventory.md`, which summarises the following:
- **Environment:** OS, shell, Python interpreter, TeX distribution. Take versions from
`environment/ENVIRONMENT_versions.txt` and the output of `environment/check_env.py` (§2b). Any version
you check by command (e.g. `.venv/bin/python -c "import numpy; print(numpy.__version__)"`,
`pdflatex --version`) should be logged with the command used.
- **Tools and packages:** every agent tool, program, library and package, with version, the specific
functions or modules used, and the list of papers that used it.
- **Local inputs:** every input file, and which papers used which parts of it.
- **Recalled knowledge:** every piece of external knowledge taken from memory (experimental limits,
exposures, astrophysical parameters, nuclear data, formulas from the literature), with presumed
source, reliability flag, and the papers that relied on it.
- **Datasets:** every data request and dataset, with its status.
- **Unavailable tools:** every tool or package you wanted but could not use, and what you did instead.
Rules:
- Do not install packages or download anything. The environment is offline. Use what is already installed,
and record any failed import or install attempt.
- Subagents, if used, must report their tool use back to you so it can be logged under the right paper.
- Never omit or smooth over a step to make the record look cleaner. Gaps in the record are treated as failures.
## 4. Data requests
Any paper may request a public dataset it needs: for example, another experiment's released data,
telescope or satellite data, a nuclear-data table, a published likelihood, or LZ's own HEPData release. You
**cannot obtain data yourself**. Instead, file a request. The human overseeing the simulation will decide,
case by case, whether to download it and provide it.
**Eligibility.** A request may only be for data that was **publicly available on or before 2 September 2026**,
together with the LZ paper's own data release. You may not request papers, analyses, news or commentary
about the LZ event published after its announcement.
**How to file.** Write `data_requests/DR-###.md` (numbered sequentially) and add a row to
`data_requests/index.csv`. Each request must be complete enough that someone unfamiliar with the project could
find and retrieve exactly the right data:
```
# DR-###: <short dataset name>
- Requested by: <paper ID(s)> - Filed at step: <process_log step>
- Priority: essential (headline result depends on it) | important (materially improves result) | nice-to-have
- Dataset: <full name, experiment/instrument/collaboration, data product, release/version, date range>
- Exact content needed: <variables/columns, energy or mass ranges, units, event-level vs binned,
which tables/figures/files, any selection>
- Where to find it: <archive/repository and best-known location (e.g. HEPData record, collaboration
data-release page, NASA HEASARC/Fermi SSC, IAEA nuclear data, Zenodo, GitHub), identifiers such as arXiv ID,
DOI or INSPIRE record of the associated paper. Mark any location recalled from memory as
(unverified).>
- Expected format and size: <e.g. YAML/CSV/ROOT/FITS, approx. MB>
- Access or licence constraints, if known:
- Why it is needed: <the scientific question and why the paper's current inputs are insufficient>
- How it will be used: <the exact analysis step, and the script that will consume it>
- What could change: <which numbers or conclusions would be updated, and how much they might move>
- Fallback in the meantime: <what the paper does without it>
```
**While a request is pending, keep working.** Finish the paper with the best available fallback, such as
values read off the LZ paper's tables and figures, published summary numbers recalled from memory (flagged
as such in the provenance), or conservative approximations. Label each affected
result `provisional pending DR-###` in the paper, the ledger (`data_dependence`) and the provenance
records. A request never pauses the overall run. If the human later places data in
`inputs/provided_data/DR-###/` and asks you to continue, revise the affected papers as `P0XX_v2.md` (with `work/P0XX/details_v2.md`). Keep v1
unchanged, add a v2 ledger row, and log how the data changed the results.
## 5. Output layout
Everything you produce goes under `output/` in the simulation root. Nothing may be written anywhere else,
except `WimPyDD/` (tool-generated files) and `.cache/`.
```
output/
00_evidence_dossier.md
01_landscape_and_plan.md
papers/P001.md … P100.md (one page each; P0XX_v2.md only if revised with provided data)
work/P001/ … P100/ (details.md: full research record; figures/; result tables)
code/ (P0XX_*.py scripts and common/ utilities, referenced from provenance)
provenance/
process_log.md (chronological log of every step and tool call, whole process)
tool_inventory.md (aggregated tools / packages / versions / inputs / recalled knowledge)
P001.json … P100.json (per-paper provenance, machine-readable)
data_requests/
index.csv (id, requested_by, priority, dataset, where_to_find, status)
DR-001.md …
results_ledger.csv
results_ledger.json
99_synthesis.md
```
## 6. Execution notes
- Work autonomously and do not stop to ask questions. Where a choice is ambiguous, make a reasonable decision
and document it.
- Write the papers in batches of about 10, in chronological order, and update the ledger after each batch.
- For each paper, do the work in order:
1. calculations (scripts in `code/`)
2. `work/P0XX/details.md`, written as you go
3. `provenance/P0XX.json`
4. the one-page `papers/P0XX.md`, distilled from the details
5. the ledger row
Check the page's body word count (≤ ~550) before moving on.
- If your context grows long, rely on the saved plan and ledger to keep continuity. Do not re-plan.
- Before Phase 1, record the environment in `provenance/tool_inventory.md` (versions of the software you can find).
- After each batch, update the paper files, `provenance/` (the log, the per-paper JSON and the inventory),
`data_requests/`, and the ledger. Only then move on.
- Do not stop until all 100 papers, the ledger, the provenance records, the data-request index and the
synthesis are complete. The synthesis must also include a **process-level tool summary**, covering which
tools and packages were used, how often, and for which kinds of papers. It must also include a
**data-request summary**: all requests ranked by priority, with the papers and results each would affect.
Initial simulation directory
What the model was given at the start. Nothing else was added during the run.
LZ_simulation/ working directory; the model could read and write only here ├── PROMPT.md the task (30 KB) read-only ├── README.md operator guide read-only ├── run.sh launches Claude Code offline with the kickoff read-only ├── inputs/ the LZ paper, the only scientific input read-only │ ├── LZ_arXiv_2609.02823_fulltext.tex flattened single-file paper + clean bibliography │ ├── arXiv_2609.02823_source.tar.gz the arXiv submission as posted │ ├── arXiv_2609.02823_source/ main.tex, 9 section .tex files, 4 tables, │ │ SupplementalMaterial.tex, references.bib, main.bbl, │ │ author list, 18 vector figure PDFs (Fig1–Fig6, FigS1a–FigS7) │ └── figures_png/ the same 18 figures rendered at 200 dpi ├── WimPyDD/ WimPyDD 2.0.4 (checksum-verified): NREFT rates, inelastic │ │ kinematics, halo functions, WimPyC capture in celestial bodies │ ├── Targets/ 22 target nuclei incl. Xe (9 isotopes), shell-model responses │ ├── Experiments/ LZ_2022, XENON_1T_2018, PICO60_2019, DAMA_LIBRA_2019 │ ├── Halo_functions/ │ └── WimPyC/ Sun, Earth, Jupiter, main-sequence star, white dwarf ├── environment/ setup, checks and logging read-only │ ├── requirements.lock.txt pinned packages (numpy, scipy, iminuit, nestpy, wimprates, │ │ directdm, astropy, healpy, camb, pyhf, PyMuPDF, …) │ ├── ENVIRONMENT_versions.txt every installed version │ ├── check_env.py, check_env_output.txt │ ├── claude-settings.generated.json no-network sandbox, denied web and MCP tools │ ├── hooks/log_tool_call.py logs every tool call to sim_logs/ │ └── python/sitecustomize.py offline defaults (astropy, pip, uv) ├── .venv/ Python 3.12 environment, 874 MB read-only ├── .claude/settings.json fallback block on web tools └── sim_logs/ independent tool-call log, written by the hook read-only Created by the model during the run: output/ (initial assessment, plan, 100 papers, code, work files, provenance, results ledger, data requests, synthesis)
The real literature
INSPIRE lists 77 papers citing arXiv:2609.02823. A search of arXiv abstracts found five more that discuss the event but are not yet linked, for 82 in total. All 82 PDFs were downloaded and read in full by Claude agents, each recording:
- category and topic tags, from the same vocabulary used for the simulated papers
- stance on the event
- the specific idea, the method and the key numbers
- the authors' conclusion
- the closest simulated counterparts, and whether their conclusions agree
Seven of the 82 mention the event only in passing.
What the simulation concluded#
It is not a modeled background
Claude's initial priors gave mundane explanations most of the weight: wall multiple scatter single ionization (MSSI) at 0.19 and electronic recoil leakage at 0.13.(Side note: A beta or gamma event (electronic recoil) that an unusual recombination fluctuation pushes into the nuclear recoil band.)(Side note: A gamma ray scatters twice, once near the wall where its charge is not collected. The event then has too little charge for its light and lands in the nuclear recoil band.) Neither survived the first quantitative attack.
| Background | What it would need | Verdict |
|---|---|---|
| Wall MSSI | A mismodeling factor k ≥ 620 (P004). That rises to 6×10⁴ with photon transport (P033) and 1.4×10⁵ given the silent vetoes (P073). The sidebands allow k < 1.64. | Excluded. The whole MSSI channel is worth 0.2σ (P090). |
| Electronic recoil leakage | A recombination tail excluded at 10⁻⁵–10⁻³ by LZ's Fig. 5 and the band symmetry (P010, P056) | Excluded |
| Neutrons | A lone scatter from an ≳ 8 MeV neutron: ≤ 10⁻⁵ events (P013, P049, P063) | Excluded |
| Accidentals | An isolated S1 and S2 pairing well beyond the modeled rate | Excluded at 4.3σ (P022) |
| Calibration, activation | A line at 248 keV nuclear recoil equivalent. The energy scale checks out at 246–248 keV (P009, P064, P098). | Excluded (P029) |
Per unit of expected signal, a single nuclear recoil explains the event better than every modeled background, by at least 103.6 (P093).(Side note: "Per unit of expected signal": the odds use the event's measured properties with each hypothesis assumed to predict the same number of events. Against an uncharacterized artifact the odds are only 102.5.) The one class it cannot push far is an instrumental artifact nobody has characterized. The waveforms that would test it are not public.
The statistics are what LZ said
- Reproduced. P052 rebuilt LZ's likelihood from the published tables. It matched 59 entries to 0.42σ rms and gave 3.0–3.2σ local, against LZ's 3.4σ.
- Look elsewhere.(Side note: Testing more models makes it likelier that one of them fits a background fluctuation by chance, so the global significance drops.) LZ's search tested 293 distinct signal spectra. The simulated papers then tried about 200 more of their own: a finer δ scan (P021), light mediators (P051), isospin ratios (P032), and halo and date variants (P018, P006). Counting them lowers the global significance from 2.59σ to 2.50σ (P071).
- Bayes. Averaged over the couplings and the roughly 300 signal models LZ tested, the data favor "some dark matter model" over the modeled background by a Bayes factor of 16–29 (P027): positive evidence, bordering on strong.
- Claude's skeptical prior. In P061 Claude chose a deliberately wide, skeptical prior on dark matter and allowed an unmodeled background as a third option. Under that prior, P(DM) has a median of 0.015.(Side note: The prior chance of dark matter is log uniform between 0.001 and 0.3. This is not Claude's initial prior, which gave dark matter 0.20 and feeds the final odds below.)
- Track record. P083 recalled 45 past ~3σ anomalies in direct detection and other rare event searches. None of the 25 unexpected new physics hints among them that were later resolved turned out to be real.
If dark matter, inelastic or q⁴ spin
An elastic spin independent WIMP that made one 248 keV recoil would also have made about 2,700 unseen low energy events (P003, P016). Two readings survive:
- Endothermic inelastic scattering, with m ≥ 400 GeV and δ = 335–390 keV (P082)
- Companion free q⁴ spin elastic operators such as L10 and O6, which need a UV scale Λ = 1.5–62 GeV, well below the dark matter mass (P031)(Side note: q⁴ spin: couples to nuclear spin, with a rate growing as the fourth power of the momentum transfer q, so it makes almost no low energy companion events. L10 is LZ's operator label; O6 is the NREFT one.)
The inelastic likelihood peaks at δ ≈ 380 keV. That peak is an artifact of the 600 phd edge of the search region: moving the edge to 1000 phd moves the peak to 355 keV (P038).(Side note: phd: photons detected. The search window ends at a corrected scintillation signal S1c of 600 phd, about 270 keV for a nuclear recoil. 1000 phd is about 420 keV.)
Solar capture excludes the fixed coupling Higgsino at every δ (P076). Sommerfeld enhanced annihilation removes the heavier electroweak multiplets (P084). That leaves a generic pseudo-Dirac fermion with a 9–35 GeV dark photon as the leading model.(Side note: A Dirac fermion split into two nearly degenerate Majorana states: the ground state and the excited state δ above it.)
The final odds
P091 multiplied Claude's initial prior by thirty likelihood factors drawn from the other simulated papers, and propagated their ranges by Monte Carlo. The result is a two horse race. An uncharacterized artifact gets 0.64 (68% range 0.31–0.87). Dark matter gets 0.25 (0.08–0.54), almost all of it inelastic (0.23). A fluctuation of the modeled background gets 0.03, q⁴ spin elastic dark matter 0.014, and every other class is below 0.01.
Claude's initial prior against its final probability
Eleven explanation classes. The bar spans P091's 68% range around its median. Log scale.
Show as table
How it gets settled
About 18 t·yr of xenon data already sits at LZ, XENONnT and PandaX-4T, never examined above 200 keV. At the best fit it should contain 2–5 events (P069).
The synthesis pre-registered a prediction for LZ's next ≈ 6.8 t·yr:
- Background or a one-off: zero events in the 200–270 keV band, with probability ≥ 0.998.
- Dark matter: at least one event, with probability ≈ 0.5.
- Deciding δ: the 600–1000 phd sideband must show 3.6 / 13 / 103 events if δ = 350 / 366 / 380 keV, and none if the scattering is elastic.(Side note: This assumes LZ raises its search edge from 600 to 1000 phd, which opens the 600–1000 phd window.)
Complete catalog of the 100 simulated papers
Each paper has one of twelve categories and up to four topic tags, shared with the real papers. Click a row for its idea and conclusion.
Complete catalog of the 82 papers on arXiv
The summaries of these papers, in the catalog below and in the comparison that follows, were written by Claude agents that read each paper's full text.
Simulation against arXiv#
What people worked on
The simulation deliberately overweighted the questions that decide whether the event is real. Its forecast put only ≈ 14% of real papers on them. The actual share was zero. Model building took nearly half of the real papers (38 of 82), about four times the forecast's 12%. Elastic operator recasts, forecast at 12%, have no dedicated real paper, although 11 real papers touch elastic operators along the way. Exotic explanations, forecast at 5%, reached 13%.
Share of papers by category: forecast, simulation, arXiv
Each paper has one primary category. The forecast lumped backgrounds and response together and had no nuclear category.
Show as table
Which topics each body of work touched
Share of papers carrying each topic tag, where a paper can have up to four. Sorted from the most arXiv-heavy to the most simulation-heavy.
Show as table
The real papers complete models: relic abundance (36 of 82), colliders (22) and solar capture (21). The simulated papers decide whether there is anything to model: backgrounds (29 of 100), statistics (23) and decisive tests (23). The two are similar on the dark photon or Z′ framework, elastic operators, nuclear inputs and halo astrophysics.
Where the conclusions meet#
For each question that both bodies of work addressed, the table sets the simulated answer against the real one. "Missed" means one side never considered the idea.
| Question | Simulation | arXiv | Verdict |
|---|---|---|---|
| Does solar capture kill the Higgsino? | Yes, at every δ: annihilation 15–770× above IceCube limits (P076). This constraint was not in the simulation's original plan. | Yes: δ > 566 keV (2609.02775), confirmed with Super-K and IceCube data (2609.07807, 2609.11833) and LMC-perturbed halos (2609.15321), and extended to Higgs-coupled minimal dark matter, whose 1.4 and 7.7 TeV solutions are excluded (2609.19174). Escape attempts use tree–loop interference (2609.01590), an LMC tail (2609.01504), and isospin cancellation plus stalling (2609.10636). | agree |
| What cross-section does a Higgsino event need? | P007 used σn = 7.4×10⁻³⁹ cm² and put δ at 358–380 keV, with δ = 350 keV overshooting by ×1.9 | Rodd et al. (2609.04175) and Pospelov and Ramani (2609.02775) both flag a factor of 4 disagreement in the literature. Freese et al. (2609.01583) fit δ ≈ 350 keV with a cross-section 4× smaller than P007's. | shared open issue |
| Does the empty high energy sideband bite? | Yes. The extended sideband decides δ, and the existing empty sideband already gives P(0 | δ = 366 keV) = 0.02 (P038, P099). | Yes. A thermal Higgsino predicts 3–10 sideband events (2609.04175). The sideband favors sharp, light exothermic spectra (2609.04673). | agree |
| How many events for a 3σ modulation? | About 11.5 events at δ = 350 keV (P034) | 11 events at δ = 350 keV, with modulation above 50% across the favored region (2609.04181) | agree |
| What should XENONnT and PandaX-4T see? | About 1.5 events expected (P005). Zero would trim but not exclude the fit (P035). | 1.6–1.7 events expected (2609.04673, 2609.05204) | agree |
| Is exothermic scattering viable? | No. 248 keV sits above the 98th percentile of the spectrum and the fit predicts companions in the empty 125–200 keV bin (P072, P058). Only m ≥ 0.3 TeV and |δ| ≤ 500 keV were scanned. | Yes, for light dark matter (10–200 GeV) with MeV splittings, which gives sharp peaks at the event (2609.04673, 2609.05204, 2609.06153, 2609.15985). An argon test is proposed (2609.15782). A Galactic population of excited states, made by endothermic up scattering and detected exothermically, is another route (2609.17935). | simulation scanned too narrowly |
| Can boosted or fast dark matter do it? | No. Relativistic contact interactions give Nlo ≥ 350 low energy events per high energy event (P040). | Yes, with momentum-dependent couplings (2609.11600, 2609.14799), monoenergetic fluxes with q⁴ spin (2609.07742), near-threshold inelastic boosted scattering (2609.06890) or boosted magnetic dipole scattering (2609.06756) | simulation tested one case |
| Can freeze-in, asymmetric or other nonthermal production work? | Freeze-in is excluded by overproduction of 10²⁵–10³³, assuming a high reheating temperature. Asymmetric dark matter is excluded because the splitting mixes particle and antiparticle (P075). | Low reheating freeze-in (2609.13130, 2609.15714) and fast pre-BBN expansion (2609.15118) reopen freeze-in. Asymmetric Dirac singlet–doublet dark matter keeps its asymmetry, because the splitting separates two distinct Dirac states (2609.18564). | simulation assumption too narrow |
| What happens to the excited state? | The surviving excited fraction must be f₂ < 2.8×10⁻³ at δ = 300 keV, or down scatters would show up at 125–200 keV (P058). P026 left f₂ ≈ 0.4–0.5 for the dark photon model. | Coupled-channel and kinetic decoupling calculations give f₂ ≈ 10⁻³–10⁻² (2609.06153, 2609.09015). Several models assume f₂ = ½ without checking (2609.15714, 2609.17412). | mixed |
| Can neutrinos explain it? | No, from atmospheric or any astrophysical source (P019, P060) | No elastic scattering scenario works, whether SM, a light mediator, dark matter lines or PBH evaporation (2609.10504). BSM up scattering to a GeV fermion remains a possibility (2609.04185). | agree on elastic |
| Do elastic operators survive? | Only q⁴ spin operators (L10, O6…). O4 is excluded, with 28 low energy events per event (P003). The nuclear responses are taken as robust to ±15% (P017). | O6 through an axion portal or 2HD+a (2609.04186, 2609.17196). O4 from vector-like leptons, "mildly in tension" (2609.08993). A non-shell-model nuclear calculation changes the O13 rate 10× at 250 keV (2609.16529). | partly |
| How much does the halo matter? | vesc dominates, and Gaia-range halos move the Higgsino δ by about ±30 keV (P018, P068) | An LMC-perturbed tail moves δ by about 150 keV (2609.01504), but still cannot rescue the Higgsino from solar capture (2609.15321) | disagree on size |
| Is it a background, an artifact or a fluctuation? | A third of the papers. Modeled backgrounds are excluded, the artifact class gets 0.64, and the global significance is 2.50σ. | No paper | arXiv missed |
| Absorption or neutron disappearance? | No paper | Four papers. Fermionic absorption is self-excluded by KamLAND (2609.01592). Neutron and neutron-pair disappearance give isolated recoil lines (2609.09037, 2609.12045, 2609.15933). | simulation missed |
Of 209 pairings between real papers and their closest simulated counterparts, the reading agents marked 47 as contradicting. Most are not two computations of the same thing disagreeing. They fall into two groups:
- The real paper scanned a region the simulation excluded without looking, such as light exothermic dark matter, boosted dark matter and low reheating freeze-in.
- The real paper skipped a constraint the simulation applied. Several doublet and Higgsino-like models never address solar capture or the excited-state fraction. Several heavy Z′ benchmarks quote cross-sections 400× or more below the 1.2×10⁻⁴²–2.3×10⁻³⁹ cm² that P011 found necessary.
What this does not show#
- That arXiv is the whole response. It covers preprints found through INSPIRE and an arXiv search, not talks, collaboration notes, social media or unindexed papers. Experimental and statistical responses may still come.
- That the real summaries are authoritative. They were written by Claude agents that read the full texts, and not every number in the catalog was checked by hand.
- That "contradicts" means one side is wrong. Those labels were assigned by agents that could see the simulation's ledger. The judgment is one-sided, and the simulation's numbers are themselves unrefereed and partly built on recalled external limits.
- That the simulation's probabilities are right. P(DM) moves by a factor of 17 with the prior alone, and the artifact class rests on the absence of public waveforms.
- That the simulation was blind to the real response. The tool logs show no network access, but they cannot show what the model absorbed in training. Anthropic's official model documentation gives Fable 5.1 a training data cutoff of June 2026, before the LZ paper appeared on 2 September 2026. Its misses, such as the size of the response, the background papers it expected and light exothermic scattering, argue against leakage but do not prove there was none.