Skip to content

Sources register

The evidence behind the skill's slow-rotting claims. Every claim the skill makes about the outside world that will not change week to week has an entry here; "where did you get this?" is answered from this file, in any phrasing (skills/advisor/SKILL.md → the laws).

What belongs here — and what never does. The register holds slow-rotting canon only: findings, methods and standards that age in years. Fast-rotting facts never enter — a price, a current API limit, a competitor's live feature stay fetch-at-decision-time rules, quoted with their check-date at the moment of use, never cached here to rot.

One fixed form per entry, so a wrong entry is visibly wrong: id · full citation · live URL/DOI/arXiv · archive link · licence · one-paragraph distillate (our words) · check-date · cited-by.

Cited-by names files, never line numbers. A line number is wrong the next time anything above it is edited, and every entry here once pointed at a line — in files that no longer exist. The durable form is the file that carries the claim; the id is greppable when the exact place is wanted.

Licence tiers (they decide what we may hold): free — an open licence (MIT, CC, a public standard); a copy may be carried later. copyrighted — citation + archive link + our own distillate, never the text itself. math — a formula recorded by name, which is not copyrightable. Each entry states its tier.

Upkeep. scripts/fetch-source.py builds and checks these entries: --resolve <doi|arxiv|url> prints a skeleton, --archive <url> triggers a Wayback snapshot, --verify walks every live URL. --verify runs at every release (./contributing → The release ritual).

Persona theatre — grounding, fidelity, and the limits of synthetic audiences

park-self-reports · Park et al., self-report-grounded individual simulation

  • Citation: Park, J.S., et al. "LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals." arXiv:2411.10109 (2024; v1 was "Generative Agent Simulations of 1,000 People").
  • Live: https://arxiv.org/abs/2411.10109
  • Archive: http://web.archive.org/web/20260726212521/https://arxiv.org/abs/2411.10109
  • Licence: copyrighted (author © under arXiv's non-exclusive distribution licence) — cite + archive + our distillate
  • Distillate: Agents built from a person's own self-reports reproduce that person's survey answers at 83% (interview-grounded) / 82% (survey-grounded) / 86% (both) of the person's two-week test-retest ceiling, versus 74% for demographics-only; a free-text "persona paragraph" scores 0.71, below even the demographics baseline (0.74). Self-report grounding also reduces accuracy disparities across racial and ideological groups. The takeaway the skill leans on: the grounding artifact — the interview transcript — is the product, not a written bio.
  • Check-date: 2026-07-27
  • Cited-by: audience.md · templates/PERSONA-template.md

park-hai-brief · Park et al., Stanford HAI policy brief

ashokkumar-nature · Ashokkumar et al., direction not magnitude

  • Citation: Ashokkumar, A., Hewitt, L., Ghezae, I., Willer, R. Nature (advance online publication, 2026-07-08). doi:10.1038/s41586-026-10742-x.
  • Live: https://doi.org/10.1038/s41586-026-10742-x
  • Archive: pending (Save Page Now triggered 2026-07-27; availability: https://archive.org/wayback/available?url=https://doi.org/10.1038/s41586-026-10742-x)
  • Licence: copyrighted (Springer Nature) — cite + archive + our distillate
  • Distillate: Across a large replication set, LLM simulations track the direction of experimental effects at about r≈0.85 while systematically overestimating their magnitude. This is the evidence for the theatre's hardest rule: a synthetic verdict may state direction, never a magnitude (no "23% would churn").
  • Check-date: 2026-07-27
  • Cited-by: audience.md · skills/advisor/SKILL.md (the second pyramid)

ls-types · Lewis & Sauro, a taxonomy of synthetic users

ls-review · Lewis & Sauro, a review of synthetic-user experiments

sharma-sycophancy · Sharma et al., sycophancy is trained in

  • Citation: Sharma, M., et al. "Towards Understanding Sycophancy in Language Models." ICLR 2024. arXiv:2310.13548 (2023).
  • Live: https://arxiv.org/abs/2310.13548
  • Archive: http://web.archive.org/web/20260725125159/https://arxiv.org/abs/2310.13548
  • Licence: copyrighted (author © under arXiv's non-exclusive distribution licence) — cite + archive + our distillate
  • Distillate: Sycophancy — telling the user what they want to hear — is a trained-in property of RLHF'd assistants, consistent across several models and tasks. A persona built on such a model inherits that compliance, which is why the theatre's calibration layer suppresses sycophancy explicitly (a synthetic respondent is a pleaser unless corrected).
  • Check-date: 2026-07-27
  • Cited-by: audience.md · skills/advisor/SKILL.md (useful over agreeable)

tjuatja-biases · Tjuatja et al., LLM response biases ≠ human ones

  • Citation: Tjuatja, L., et al. "Do LLMs Exhibit Human-like Response Biases? A Case Study in Survey Design." TACL 12 (2024). arXiv:2311.04076 (2023).
  • Live: https://arxiv.org/abs/2311.04076
  • Archive: http://web.archive.org/web/20260116061015/https://arxiv.org/abs/2311.04076
  • Licence: copyrighted (author © under arXiv's non-exclusive distribution licence) — cite + archive + our distillate
  • Distillate: Tests whether LLMs reproduce known human survey response biases (acquiescence, question-order effects) and finds their biases do not reliably mirror human ones — sometimes absent, sometimes inverted. Caveats how far a synthetic survey respondent can stand in for a human one.
  • Check-date: 2026-07-27
  • Cited-by: audience.md

argyle-silicon · Argyle et al., silicon sampling and its diversity limits

  • Citation: Argyle, L.P., et al. "Out of One, Many: Using Language Models to Simulate Human Samples." Political Analysis (2023). doi:10.1017/pan.2023.2.
  • Live: https://doi.org/10.1017/pan.2023.2
  • Archive: pending (Save Page Now blocked HTTP 523 on 2026-07-27; availability: https://archive.org/wayback/available?url=https://doi.org/10.1017/pan.2023.2)
  • Licence: copyrighted (Cambridge University Press) — cite + archive + our distillate
  • Distillate: Introduces "silicon sampling" — conditioning an LLM on demographic backstories to simulate human survey samples — and shows it can reproduce some subgroup patterns while collapsing within-group diversity. Backs the caution that synthetic samples flatten variety rather than represent it.
  • Check-date: 2026-07-27
  • Cited-by: audience.md

wang-flattening · Wang et al., identity flattening

  • Citation: Wang, A., Morgenstern, J., Dickerson, J.P. "Large language models that replace human participants can harmfully misportray and flatten identity groups." Nature Machine Intelligence (2025). arXiv:2402.01908.
  • Live: https://arxiv.org/abs/2402.01908
  • Archive: http://web.archive.org/web/20260607174738/https://arxiv.org/abs/2402.01908
  • Licence: copyrighted (author © under arXiv's non-exclusive distribution licence; journal © Springer Nature) — cite + archive + our distillate
  • Distillate: Finds that using LLMs to replace human participants can harmfully misportray and flatten identity groups — reproducing majority stereotypes and erasing within-group variation. This is the direct evidence for the never-assign-a-bias-from-demographics rule: a demographic backstory produces a caricature, not a person.
  • Check-date: 2026-07-27
  • Cited-by: audience.md

kapania-simulacrum · Kapania et al., LLMs as qualitative participants

  • Citation: Kapania, S., et al. "'Simulacrum of Stories': Examining Large Language Models as Qualitative Research Participants." CHI 2025. arXiv:2409.19430 (2024).
  • Live: https://arxiv.org/abs/2409.19430
  • Archive: http://web.archive.org/web/20260411143601/https://arxiv.org/abs/2409.19430
  • Licence: copyrighted (author © under arXiv's non-exclusive distribution licence) — cite + archive + our distillate
  • Distillate: Treating LLMs as qualitative research participants yields plausible but hollow "simulacra of stories" that miss the lived specificity of real interviews. Marks the boundary of synthetic personas in qualitative work — a supplement, never a replacement for a real transcript.
  • Check-date: 2026-07-27
  • Cited-by: audience.md · templates/PERSONA-template.md

Cost routing — cheap-first, conditional on a good verifier

frugalgpt · Chen, Zaharia & Zou, FrugalGPT

  • Citation: Chen, L., Zaharia, M., Zou, J. "FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance." arXiv:2305.05176 (2023).
  • Live: https://arxiv.org/abs/2305.05176
  • Archive: http://web.archive.org/web/20260722202256/https://arxiv.org/abs/2305.05176
  • Licence: copyrighted (author © under arXiv's non-exclusive distribution licence) — cite + archive + our distillate
  • Distillate: A cascade that queries cheaper models first and escalates only on low confidence can match the best single model's accuracy at up to −98% cost. The evidence for cheap-first-then-escalate routing at decomposition.
  • Check-date: 2026-07-27
  • Cited-by: cost.md · dispatching.md

routerbench · Hu et al., RouterBench

  • Citation: Hu, Q.J., et al. "RouterBench: A Benchmark for Multi-LLM Routing Systems." arXiv:2403.12031 (2024).
  • Live: https://arxiv.org/abs/2403.12031
  • Archive: http://web.archive.org/web/20260606022825/https://arxiv.org/abs/2403.12031
  • Licence: copyrighted (author © under arXiv's non-exclusive distribution licence) — cite + archive + our distillate
  • Distillate: Cascades beat both any individual LLM and a zero-cost router only when the verifier is good — judge error ≤0.1, deteriorating past 0.2. The load-bearing caveat: cheap-first routing is conditional on a good verifier. In the skill, the review gates are that verifier, so the condition is already met — the caveat reads as a strength, not a risk.
  • Check-date: 2026-07-27
  • Cited-by: cost.md · dispatching.md

Repository context files

agentsmd-eth · Gloaguen et al., do ./contributing files help?

  • Citation: Gloaguen, T., Mündler, N., Müller, M.N., Raychev, V., Vechev, M. (ETH Zurich). "Evaluating ./contributing: Are Repository-Level Context Files Helpful for Coding Agents?" arXiv:2602.11988 (v2, 2026-06-23).
  • Live: https://arxiv.org/abs/2602.11988
  • Archive: http://web.archive.org/web/20260711121106/https://arxiv.org/abs/2602.11988
  • Licence: copyrighted (author © under arXiv's non-exclusive distribution licence) — cite + archive + our distillate
  • Distillate: A coding-agent benchmark finds repository-level context files (./contributing) do not improve task success rates and add roughly +20% inference cost; LLM-generated context files perform slightly worse than none. The evidence behind "curate the shared guide, don't autogenerate it."
  • Check-date: 2026-07-27
  • Cited-by: project-layout.md · AGENTS.md

Method provenance and standards (references, not evidence claims)

method-provenance · adapted methods, nothing embedded

  • Citation: cookiy — user-research-skill (MIT). agentman — "Synthetic Persona Creator" skill (concepts only).
  • Live: https://github.com/cookiy-ai/user-research-skill · https://agentman.ai/agentskills/skill/synthetic-persona-creator (owner-confirmed 2026-07-27; page verified live: three-layer persona architecture — identity foundation · context seeding · response calibration — plus cohort distribution)
  • Archive: http://web.archive.org/web/20260411081018/https://github.com/cookiy-ai/user-research-skill
  • Licence: free — cookiy's user-research-skill is MIT (a copy may be carried later; today only the method shape is adapted, nothing embedded). agentman: no licence stated on the page; concepts taken, nothing embedded (no code carried, so no licence obligation).
  • Distillate: Method lineage, not evidence. cookiy's MIT skill supplied the shape of the qualitative-research flows (its qualitative-research-planner → our persona-interview flow, its synthesize-research-report → our QDA step), adapted through the import gate. agentman supplied the calibration and cohort concepts behind the persona response-calibration layer. Recorded so every adaptation is auditable and no vendor wrapper is smuggled in.
  • Check-date: 2026-07-27
  • Cited-by: TRADEMARKS.md · PATTERNS.md

standards-cluster · named review standards