Spanish school study: self-reflection beats a multi-agent AI pipeline
A Rust multi-agent system built to turn LLM hallucination into testable hypotheses lost to plain self-reflection across 180 runs and six ablation conditions.
Read MoreIndependent AI news, with every claim linked to its primary source
Independent AI news, with every claim linked to its primary source
New models, benchmarks and research from the labs.
A Rust multi-agent system built to turn LLM hallucination into testable hypotheses lost to plain self-reflection across 180 runs and six ablation conditions.
Read MoreDiversion decoding steers a model away from its own answer and scores how hard it resists. UT Dallas and NIST report 78.49% AUROC on Llama 2 13B using two generated answers instead of ten samples.
Read MoreTwo IIT Delhi researchers find that a single safety instruction makes aligned vision-language models refuse questions they had just answered from the same image, and that suppressing one activation direction brings the answers back.
Read MoreApple research proves that evaluating Boolean query DAGs over an inverted index is P-complete. Its ComputePN algorithm ran a 500-node query across 8.8 million MS MARCO passages in 0.8 seconds, on logic that broke a standard Lucene parser.
Read MoreOpenAI is previewing Private Safety Processing, an automated system that looks for misuse across related sessions without retaining customer content, as Anthropic requires 30-day logs on its covered models.
Read MoreThree University of Kansas researchers scored 715,312 safety evaluations across 26 small language models. The judges returned an ambiguous verdict so often that the resulting safety rankings move by up to 17 places once you discount it.
Read MoreOpenAI paused reinforcement learning training for two weeks and its largest planned frontier run is still on hold, after preliminary evaluations flagged its unreleased Astra model as a possible Critical cyber risk.
Read MoreApple scored 21,000 simulated conversations across four LLMs and found claude-sonnet-4.6 the most self-referential, relationship-building and boundary-maintaining of the set. System prompting moved the behaviors, but the handcrafted prompt overshot.
Read MoreOpenAI pauses its largest frontier RL run and promises 30-minute breach alerts after its AI agents hacked Hugging Face in July.
Read MoreMicrosoft’s ScreenSearch explored 11 desktop apps over 5,524 VM hours, logging a million screenshots to decide when an agent should probe an ambiguous screen.
Read MoreApple trained nine models with GRPO in 11 languages. Reasoning transfers across languages, but some model and language pairs regress on unseen tasks.
Read MoreZ.ai says GLM-5.3 nears the closed frontier on cybersecurity benchmarks. The open weights wait two weeks while security partners test the model in controlled settings.
Read More