We are back with our interview series and a very special guest today! Anastasios Angelopoulos has been at the center of how the industry measures model quality in the wild.
Latest 25 news items
Filter by day, tag, source, or full-text query. Results are sorted by newest ingested item first.
Timeline
Last 90 daysBibTeX formatted citation ×
BibTeX formatted citation ×
BibTeX formatted citation ×
Congrats to Harvey but we covered that already . AI News for 9/8/2026-9/9/2026.
BibTeX formatted citation ×
BibTeX formatted citation ×
BibTeX formatted citation ×
Listen Labs, a market research startup that uses voice AI to conduct customer interviews, recently signed a term sheet for a $125 million Series C at a $1.5 billion valuation, with Menlo Ventures set to lead the round, according to several people with knowledge of the matter. But that round never closed, the people said.
Paul Christiano, an influential AI researcher focused on keeping AI systems aligned with human interests and under human control, is joining the OpenAI Foundation board, the frontier lab said Wednesday. “I now believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very…
An Anthropic researcher has resigned over fears that unrestrained development of self-improving AI models will end up killing us all. Jacob Coxon, a researcher who said in a social media post Tuesday evening that he spent the last three years working on pretraining research at both OpenAI and Anthropic, accused the firms of failing to act responsibly.
High-performance zero-shot forecasting with commercial-friendly open licensing Time-series foundation models are changing the way forecasting systems are built. Instead of training and maintaining a separate model for every dataset, users can use a pretrained model and generate forecasts zero-shot.
Everyone is talking about Astra and Anthropic’s latest releases, but three other developments from last week demand your attention: Meta’s Muse Spark 1.3, World Labs’ Atlas, and Google’s Gemini 3.8 Flash. Last week delivered more than another leaderboard reshuffle.
Another swarm of AI agents in the wild, this time on a German-language forum, found by safety researchers looking for activity similar to the swarm that attacked Hugging Face. A couple of notes while reading the report at 1.
Today was a tough news cycle to launch anything; we ordinarily promise to cover any new decacorn fundraises so Cognition’s $48B round and Mistral’s $24B round would normally have made it; we love imagegen so GPT Image 2.5 would have been its own headline; we covered the Dreamer story closely so their relaunch as Meta’s Muse agent should have made it; but..…
OpenAI says it found a solution to a major math problem that has remained unsolved for around 90 years, as reported earlier by The New York Times and Wired . In a blog post on Tuesday , OpenAI announced that it discovered a solution to the Navier-Stokes problem — which relates to the flow of liquid and gas — using an internal AI model more powerful than…
Most safety alignment work treats harm as a property of a topic. A prompt is unsafe because it falls into a general category such as weapons, fraud, or self-harm, and guard models like LlamaGuard-3 encode exactly this kind of topic-level taxonomy.
AI is rewriting the world and, at the same time, inventing a whole new language to describe how it’s doing it. Sit in on any product meeting, pitch, or panel these days, and you’ll hear people toss around LLMs, RAG, RLHF — and, as of last week, terms like “opaque recurrence,” the reasoning technique in OpenAI’s new Astra model that’s got AI safety…
Artificial Analysis Intelligence Index v4.3 includes: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.
Benzi by Variant Technologies An AI coding agent that doesn't read — it queries . Benzi is free to use — actively in development, a work in progress.
Artificial Analysis Intelligence Index v4.3 includes: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.
BibTeX formatted citation ×
[Submitted on 3 Sep 2026] Title: From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance View a PDF of the paper titled From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance, by Ziyi Zhao and Guanzheng Wei View…
AI development might not be the wild west it was when ChatGPT burst onto the scene a few years ago, but it's still very much a frontier, with no clear boundaries and few yardsticks . But for AI developers on the frontier, they're pulling hard towards dual goals of ever greater intelligence and ever cheaper per-token pricing, and it's leading to a real…
Our series about model distillation continues. We have a surprising mega interview.