The series draws on these external references.
- Anthropic — Alignment Team Research — Anthropic
- NIST AI Risk Management Framework — NIST
- OECD AI Principles — OECD
- Sycophancy in language models (RLHF reward hacking) — Anthropic / arXiv
- On the Dangers of Stochastic Parrots — Bender, Gebru, McMillan-Major, Shmitchell — FAccT
- EU AI Act — Vulnerable Group Protections (Art. 5) — European Commission
- NIST AI RMF — Human-AI Configuration — NIST
- Replika and the rise of AI companions (research overview) — MIT Technology Review
- APA — AI in mental health practice — American Psychological Association
- TruthfulQA: Measuring How Models Mimic Human Falsehoods — Lin, Hilton, Evans — arXiv
- OpenAI — Calibrated Uncertainty Research — OpenAI
- Anthropic — Honesty in language models — Anthropic
- Deep Habits and the Cognitive Cost of Convenience — Cal Newport
- Anthropic — Constitutional AI — Anthropic
- Anthropic — Automated Alignment Researchers (Scalable Oversight) — Anthropic
- METR — Evaluating Models for Dangerous Capabilities — METR
- ARC Evals (now METR) — Model Evaluation Reports — ARC Evals
- Anthropic — Responsible Scaling Policy — Anthropic
- NIST AI RMF — Govern function — NIST
- EU AI Act — High-risk system obligations — European Commission
- Andy Grove — High Output Management — Penguin Random House
- Anthropic — Agent Safety Research — Anthropic
- UN — Rights of the Child in the Digital Environment (General Comment 25) — United Nations
- AI and Future Generations — Long-term governance — Future of Life Institute
- Jonathan Haidt — The Anxious Generation — Penguin Random House
- Virginia Eubanks — Automating Inequality — St. Martin’s Press
- Safiya Umoja Noble — Algorithms of Oppression — NYU Press
- AI Now Institute — Reports on Public-Sector AI — AI Now
- Brookings — Algorithmic bias detection and mitigation — Brookings Institution
- Wendell Berry — The Unsettling of America — Counterpoint Press
- UNESCO Recommendation on the Ethics of AI — UNESCO