AI Safety
Alignment, interpretability, regulation, and responsible AI development.
✦ Prompts
Full library →◈ News
Full timeline →China Launches World AI Cooperation Organization With 29 Founding Nations
Twenty-nine countries — including Brazil, Russia, Indonesia, Pakistan, Kazakhstan, Belarus, Serbia, Cuba, and Venezuela — signed an agreement in Shanghai on July 16 establishing the World AI Cooperation Organization (WAICO), a China-headquartered body positioned as an alternative to Western-led AI governance frameworks. Chinese Foreign Minister Wang Yi and founding-member representatives signed the accord in the presence of UN Secretary-General António Guterres. China committed to provide 5,000 AI training slots for developing countries over five years and to establish cooperation centers with ASEAN, the African Union, CELAC, the SCO, and BRICS.
OpenAI Makes GPT-5.6 Generally Available to All Users
OpenAI made the GPT-5.6 family — Sol, Terra, and Luna — generally available across ChatGPT, Codex, and the API on July 9, ending a two-week limited preview that began June 26 under a government-coordinated review. Pricing is $5/$30 (Sol), $2.50/$15 (Terra), and $1/$6 (Luna) per million input/output tokens; it is the first time since the June 12 Fable 5 export-control episode that every major US frontier lab has a publicly available flagship model simultaneously.
Fed Chair Warsh Names Marc Andreessen to Co-Lead New AI and Productivity Task Force
Federal Reserve Chair Kevin Warsh announced five new external task forces on July 9, naming a16z co-founder Marc Andreessen to co-lead the one studying AI's economic impact, alongside Stanford economist Charles I. Jones and Microsoft Xbox CEO Asha Sharma. The task force's mandate is to assess the economic impact of general-purpose technologies including AI to inform Fed policy judgments, with recommendations due by the end of 2026.
China Weighs Restricting Foreign Access to Its Most Advanced AI Models
China's Ministry of Commerce has held talks with Alibaba, ByteDance, and Z.ai about restricting overseas access to their top-tier models — including unreleased ones — with options ranging from a bar on public release to domestic-use-only limits, and possibly classifying unauthorized model disclosure as a national-security violation. No decision has been made; the talks follow Alibaba's internal ban on employee use of Claude Code.
Future of Life Institute's Summer 2026 AI Safety Index: Top Grade Is a C+
The Future of Life Institute published its Summer 2026 AI Safety Index on July 7, grading nine leading AI companies across 37 indicators in six domains on evidence collected through June 3. Anthropic took the top overall grade at C+ (2.66), leading five of six domains; OpenAI followed at C (2.28) and Google DeepMind at C (2.01); Meta landed at D+, with xAI, DeepSeek, and Mistral receiving failing grades. Reviewers noted that several labs, including Anthropic, OpenAI, Google DeepMind, and Meta, have weakened or eliminated earlier commitments to pause development if systems approached specified danger thresholds.
UN Opens First Global Dialogue on AI Governance With 169 Countries in Geneva
The inaugural UN Global Dialogue on AI Governance opened in Geneva on July 6, convening delegates from 169 countries in the most significant multilateral AI governance meeting yet held. The dialogue's independent scientific panel warned that governance safeguards are not keeping pace with capability advances and that the window for effective global coordination may be closing.
⬡ Tools
All tools →△ Concepts
All concepts →The field working to ensure AI systems do what humans actually want — now and as they become more capable.
Adversarial testing of AI systems to find failure modes before deployment.