AI news feed

Latest 25 news items

Filter by day, tag, source, or full-text query. Results are sorted by newest ingested item first.

Results
25
Showing newest 25 of 736
Clear

Timeline

Last 90 days
Tag: benchmark
Latent Spacelatent.spaceAdded
[AINews] not much happened today

Congrats to Harvey but we covered that already . AI News for 9/8/2026-9/9/2026.

benchmarkmodelpaperrelease
TechCrunch — AItechcrunch.comAdded
AI research startup Listen Labs scrubbed a $1.5B funding round for Salesforce talks

Listen Labs, a market research startup that uses voice AI to conduct customer interviews, recently signed a term sheet for a $125 million Series C at a $1.5 billion valuation, with Menlo Ventures set to lead the round, according to several people with knowledge of the matter. But that round never closed, the people said.

benchmarktool
TechCrunch — AItechcrunch.comAdded
OpenAI adds a prominent AI doomer to its board of directors

Paul Christiano, an influential AI researcher focused on keeping AI systems aligned with human interests and under human control, is joining the OpenAI Foundation board, the frontier lab said Wednesday. “I now believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very…

benchmarkmodelreleasesafety
TechCrunch — AItechcrunch.comAdded
'Gambling with our lives': Anthropic researcher quits, warns against self-improving AI

An Anthropic researcher has resigned over fears that unrestrained development of self-improving AI models will end up killing us all. Jacob Coxon, a researcher who said in a social media post Tuesday evening that he spent the last three years working on pretraining research at both OpenAI and Anthropic, accused the firms of failing to act responsibly.

benchmarkmodelreleasesafety
Hugging Face Bloghuggingface.coAdded
IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license

High-performance zero-shot forecasting with commercial-friendly open licensing Time-series foundation models are changing the way forecasting systems are built. Instead of training and maintaining a separate model for every dataset, users can use a pretrained model and generate forecasts zero-shot.

benchmarkmodelreleasetool
Latent Spacelatent.spaceAdded
[AINews] OpenAI reports Navier-Stokes singularity find in 88 hours using Astra-next, roughly 10,000 agents and 130B tokens (>$40M), a contender for second ever Millennium Prize awarded

Today was a tough news cycle to launch anything; we ordinarily promise to cover any new decacorn fundraises so Cognition’s $48B round and Mistral’s $24B round would normally have made it; we love imagegen so GPT Image 2.5 would have been its own headline; we covered the Dreamer story closely so their relaunch as Meta’s Muse agent should have made it; but..…

benchmarkmodelpaperrelease
The Verge — AItheverge.comAdded
Drama swirls around OpenAI’s legendary mathematical milestone

OpenAI says it found a solution to a major math problem that has remained unsolved for around 90 years, as reported earlier by The New York Times and Wired . In a blog post on Tuesday , OpenAI announced that it discovered a solution to the Navier-Stokes problem — which relates to the flow of liquid and gas — using an internal AI model more powerful than…

benchmarkmodelreleasetool
Hugging Face Bloghuggingface.coAdded
Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

Most safety alignment work treats harm as a property of a topic. A prompt is unsafe because it falls into a general category such as weapons, fraud, or self-harm, and guard models like LlamaGuard-3 encode exactly this kind of topic-level taxonomy.

benchmarkmodelsafetytool
TechCrunch — AItechcrunch.comAdded
Opaque recurrence, and other AI terms that you should probably know

AI is rewriting the world and, at the same time, inventing a whole new language to describe how it’s doing it. Sit in on any product meeting, pitch, or panel these days, and you’ll hear people toss around LLMs, RAG, RLHF — and, as of last week, terms like “opaque recurrence,” the reasoning technique in OpenAI’s new Astra model that’s got AI safety…

benchmarkmodelpaperrelease
Hacker Newsartificialanalysis.aiAdded
Announcing the Artificial Analysis Intelligence Index v4.3

Artificial Analysis Intelligence Index v4.3 includes: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

benchmark
Hacker Newsartificialanalysis.aiAdded
AI Model & API Providers Analysis

Artificial Analysis Intelligence Index v4.3 includes: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them.

benchmarkmodeltool
arXiv cs.AIarxiv.orgAdded
A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance

[Submitted on 3 Sep 2026] Title: From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance View a PDF of the paper titled From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance, by Ziyi Zhao and Guanzheng Wei View…

benchmarkmodeltool
Hacker Newstomshardware.comAdded
Frontier AI faces pricing reckoning as token volume explodes 25-fold - mid-tier models deliver 90% of flagship capability at one-sixth the cost

AI development might not be the wild west it was when ChatGPT burst onto the scene a few years ago, but it's still very much a frontier, with no clear boundaries and few yardsticks . But for AI developers on the frontier, they're pulling hard towards dual goals of ever greater intelligence and ever cheaper per-token pricing, and it's leading to a real…

benchmarkmodeltool