Stanford: robots score 89% in simulation and 12% in a real house
The distance between those two numbers is the most honest thing published about AI progress this year.
Read MoreIndependent AI news, with every claim linked to its primary source
Independent AI news, with every claim linked to its primary source
The distance between those two numbers is the most honest thing published about AI progress this year.
Read MoreWhen six labs are within 5% of each other, the question stops being which model is best and becomes which one you can afford to run.
Read MoreSWE-bench Verified went from 60% to near 100% in a year. A test everyone passes measures nothing.
Read MoreCompute spent thinking can substitute for compute spent training. That trade reorganised the field, and moved the bill onto you.
Read MoreAgents that work for five steps fall apart over fifty. The research says the problem is what they are carrying, not how clever they are.
Read MorePublishing weights used to mean you had lost the frontier. Now it means you want the distribution.
Read MoreDaybreak now runs Blue and Red tiers behind approval. The gate says more about the capability than the launch does.
Read MoreGoogle says its new entry-level model beat Anthropic and OpenAI across nine benchmarks. The cadence matters more than the scores.
Read MoreYou can run a complete AI stack on hardware you control. Whether you should depends on what an hour of maintenance is worth to you.
Read MoreAn effect concentrated in one part of the workforce can be invisible in a national average and severe for the people inside it.
Read MoreThe mechanism is simpler than the capability suggests, and unusually predictive of where these models fail.
Read MoreWe specify a training process, not a design. Nobody writes down what the finished network does, which means nobody starts out knowing.
Read More