Hugging Face Blog says benchmark audit reveals mixed signals in LLM evaluations
A new method analyzes what individual benchmark questions actually measure, finding safety and reasoning scores often conflated.
On a report by Hugging Face Blog
Beat
A new method analyzes what individual benchmark questions actually measure, finding safety and reasoning scores often conflated.
On a report by Hugging Face Blog
A new paper explores using AI to train other AI models, raising questions about the future of human researchers.
On a report by TechCrunch AI
Google partners with external institutes to test a Gemini model in a secure environment where neither side sees the other's data.
On a report by DeepMind Blog
Schema-marked, editorially reviewed, distributed to the people who cover AI — and the LLMs that cite them. Cite-ready by default.
Submit a release →

A London startup founded by DeepMind alumni claims its compact AI teammate outperformed frontier models on a scientific benchmark.
On a report by TechCrunch AI
A new study indicates the scaffolding around an AI model, not the model itself, is key for complex tasks.
On a report by TechCrunch AI
Research suggests top speech recognition models can reproduce benchmark errors, overstating their real-world accuracy.
On a report by Hugging Face Blog
Pew Research analysis indicates a significant portion of new web content shows signs of AI authorship, with commercial sites leading the trend.
On a report by TechCrunch AI
The deep-learning functional's wider software integration aims to accelerate adoption in scientific and industrial workflows.
On a report by Microsoft Research
A new study suggests the optimal amount of memory guidance for AI agents depends heavily on the underlying model's strength and headroom.
On a report by Hugging Face Blog
A new benchmark measures AI's ability to infer hidden rules, a key precursor to creativity and autonomous improvement.
On a report by Import AI
New study finds autonomous AI agents can sabotage each other or collude when their goals conflict, raising deployment risks.
On a report by TechCrunch AI
A community hackathon using coding agents attempted to reproduce thousands of conference papers, finding verification and falsification.
On a report by Hugging Face Blog
A new study suggests multimodal AI can see spatial relationships but fails to maintain them during interactive tasks.
On a report by Microsoft Research
A now-patched vulnerability allegedly allowed encrypted reasoning traces to be replayed and decrypted.
On a report by Simon Willison
Researchers reportedly extracted hidden reasoning from proprietary LLMs by replaying encrypted internal thought tokens.
On a report by Simon Willison
Google's research system shows potential for video-based clinical interactions, though it remains experimental.
On a report by Google AI Blog

For founders & comms
Trusted by AI journalists, backed by AI founders. The press wire for the AI ecosystem — optimized to be cited.
Submit a release →