Reasoning
Chain-of-thought, math, logic, and deep analytical thinking in AI systems.
✦ Prompts
Full library →◈ News
Full timeline →Claude Sonnet 5 Beats Sonnet 4.6 Across Benchmarks and Narrows the Gap to Opus 4.8
Coverage from The New Stack and others noted Sonnet 5 improves on Sonnet 4.6 on every published benchmark — roughly 63.2% on SWE-bench Pro (vs 58.1%), 76.1% on Terminal-bench (vs 55.4%), and 81.2% on OSWorld-Verified — and on the GDPval-AA v2 knowledge-work eval it scored 1,618, edging Opus 4.8's 1,615. Anthropic's system card also reports lower hallucination and sycophancy, better refusal of malicious requests, and stronger resistance to prompt injection in agentic contexts.
CLI-Universe Releases 6,000 Verified Terminal-Agent Trajectories
Researchers introduced CLI-Universe, a synthesis and verification pipeline for executable terminal-agent tasks. Fine-tuning Qwen3-32B on its 6,000-trajectory dataset achieved 33.4% on Terminal-Bench 2.0, reported as a new open-data result for models at or below 32 billion parameters.
NVIDIA Blackwell Sweeps Every MLPerf Training 6.0 Benchmark
NVIDIA's Blackwell platform topped every category in MLPerf Training v6.0, posting the fastest time-to-train across all benchmarks and the largest-scale submission to date — 8,192 GPUs on GB200 NVL72 systems. New GB300 NVL72 systems delivered up to 1.6x faster training than GB200 at the same rack scale, and NVIDIA set records on two new MoE pretraining workloads added this round, DeepSeek-V3 671B and GPT-OSS-20B.
OpenAI Releases LifeSciBench, a 750-Task Life-Science Research Benchmark
OpenAI published LifeSciBench on June 17, 2026, a 750-task benchmark built with 173 PhD-level scientists, grading models on free-response life-science research tasks — including genomic sequence files, chemical structures, and experimental figures — against expert-written rubrics averaging 25 criteria each (19,020 criteria total). The strongest model passed only 36.1% of all tasks and just 28.1% of artifact-heavy tasks, identifying scientific-artifact processing as the primary bottleneck for current AI systems.
Google Confirms Gemini 3.5 Pro Imminent: 2 Million Token Context and 'Deep Think' Reasoning Mode
Google confirmed Gemini 3.5 Pro is nearing launch after months of internal use and limited enterprise preview. The model features a 2-million-token context window and a new 'Deep Think' reasoning mode for complex multi-step tasks. CEO Sundar Pichai said 'give us until next month' at Google I/O on May 19; Polymarket prediction markets through early June placed mid-to-late June as the most likely launch window. Expected pricing is approximately $15/M input tokens and $60/M output tokens.
Gemini 3.5 Pro Enters Limited Vertex Preview with 2M-Token Context and Deep Think Reasoning
Google opened Gemini 3.5 Pro to select Vertex AI enterprise customers in limited preview in early June 2026, ahead of a broader GA Sundar Pichai described as 'next month' at Google I/O. The model targets a 2-million-token context window, Deep Think multi-step reasoning, and frontier multimodal understanding — absorbing the use cases previously routed to Gemini Ultra.
⬡ Tools
All tools →△ Concepts
All concepts →AI trained on vast text to understand and generate language.
Letting an AI invoke real code and APIs mid-reasoning.
Prompting a model to reason step-by-step before giving its final answer.
Standardised tests that measure what a model can actually do.
AI that thinks through a problem step by step before giving its final answer.