Skip to main content
Zeitgeist — a spike by Chris Gathercole
  1. Creators/

Arvind Narayanan & Sayash Kapoor — AI as Normal Technology (formerly AI Snake Oil)

About #

Arvind Narayanan (Princeton CS professor, director of the Center for Information Technology Policy) and Sayash Kapoor (Princeton CITP researcher) write empirically-grounded critiques of AI capability claims and vendor hype. Authors of the 2024 book “AI Snake Oil”; their newsletter was renamed “AI as Normal Technology” in 2025 to reflect a shift toward analyzing AI as transformative-but-ordinary technology rather than existential risk or hype.


2026-04-16 — Open-world evaluations for measuring frontier AI capabilities #

Newsletter · Read

  • Traditional benchmarks (SWE-Bench, ARC-AGI) saturate within ~2 years and are gameable via RL optimization, so the authors propose “open-world evaluations”: real-world, multi-day tasks with small samples, permitted human intervention, and qualitative log review instead of a single metric.
  • Their CRUX project had an agent attempt to autonomously publish an iOS app; it came extremely close, needing only one manual intervention — and fabricated a phone number that went undetected by App Store reviewers.
  • Argue these evaluations give institutions (e.g., app stores) early warning of emergent autonomous capabilities before they cause harm at scale.
  • Practical takeaway: evaluators should publish costs, run dry-runs, and release full agent transcripts so results are independently verifiable rather than vendor-narrated.

2026-05-21 — Do AI Risks Require Extraordinary Government Intervention? #

Newsletter · Read

  • Argue against extraordinary interventions like AI nonproliferation or strict release restrictions, since AI techniques (unlike nuclear material) are widely known and replicable within months — enforcement would be unsustainable.
  • Historical precedent (encryption controls, bomb-making info restrictions) shows such restrictions tend to expand scope over time while burdening legitimate companies more than bad actors.
  • Point to cybersecurity as the working model: vulnerability-detection tools are freely available, and defense improved via bug bounties, hardened OS/browsers, and automated testing rather than access restriction.
  • Practical takeaway: invest in distributed societal resilience — red-teaming for hospitals/schools/power grids, biosecurity material screening — rather than centralized control of AI development.

2026-05-22 — Did Google’s AI agents really build an operating system for $916? #

Newsletter · Read

  • Critique Google’s claim that AI agents built an OS from “a single prompt” for $916.92 — the prompt was later revealed to be thousands of lines long, and the number of failed attempts before the successful run was never disclosed.
  • No release of the actual prompt, source code, or execution logs, and no analysis of whether the agent produced original code versus reproducing known toy-OS implementations from training data (a common undergrad project).
  • Credit Google for disclosing cost and token usage (2.6B tokens) — more transparency than most vendor capability claims offer.
  • Practical takeaway: call for shared methodological norms for “open-world evaluations” and for independent academic/nonprofit/government verification rather than relying on vendor-generated narratives.

2026-06-11 — Why AI hasn’t replaced software engineers, and won’t #

Newsletter · Read

  • Argue mass AI-driven layoffs of software engineers aren’t happening and won’t, treating this as demonstrable evidence rather than speculation since coding is AI’s most advanced domain.
  • Cite “AI washing”: 59% of hiring managers admit exaggerating AI’s role in layoffs for optics, and New York’s WARN Act AI-disclosure checkbox was checked by essentially no filers in its first year (~0.2% of job losses).
  • Introduce the “decide-execute-deliver sandwich”: AI compresses the execution (coding) layer, but decision and delivery/verification remain human bottlenecks — a Fed study found AI increased code output 8x but releases only 30%.
  • Practical takeaway: expect the engineer role to shift toward supervising AI agents (“crane operators”) rather than disappearing — cheaper software creation likely increases overall demand (Jevons’ paradox) even as individual impact varies by seniority and firm.

2026-07-09 — Up the Stack: How AI’s Escape From the Commodity Trap Risks Enterprise Lock-in #

Newsletter · Read

  • Argue frontier AI labs face a “commodity trap” in inference pricing (undifferentiated models, low switching costs, similar capital costs) that pushes prices toward marginal cost per the Bertrand paradox, echoing how railroads and telecom infrastructure builders failed to capture the value they created.
  • Labs are escaping by moving up the stack: building embedding moats (customer data/workflows), developer ecosystems, multi-year commercial lock-in, and outcome-based pricing — evidenced by ChatGPT’s revenue dominance over the API and Claude Code’s growth.
  • Note cloud computing and TSMC as historical partial exceptions that escaped commoditization via software-like properties or near-monopoly.
  • Practical takeaway: urge competition regulators to act early — establishing interoperability standards and switching-cost transparency before lock-in compounds — and urge enterprises to evaluate portability before committing.

2026-07-13 — What will be left for us to work on? #

Newsletter · Read

  • Arvind Narayanan’s ICML 2026 keynote (Seoul): the “AI as Normal Technology” framework still holds for medium-term impact absent a discontinuity like recursive self-improvement, which warrants concern but has no imminent lab milestone that would eliminate human employment overnight.
  • Presents Princeton research showing a capability-reliability gap in AI agents: accuracy improved dramatically over 24 months, but reliability (consistency, robustness, calibration, operational safety) rose only 5-10 percentage points.
  • Distinguishes “automation agents” (need high reliability) from “collaboration agents” (tolerate unpredictability); argues collaboration agents will keep winning near-term since general-purpose + high-stakes + full automation can’t all hold at once.
  • Draws the electricity-adoption parallel: factory electrification took ~40 years of organizational reinvention, not just swapping generators for steam engines — jobs will transform radically but not vanish overnight, moving toward human-AI “co-superintelligence.”