Zo Computer:
OpenAI released GPT-6 Astra, what the AI industry has been waiting for a long time.
There are many benchmarks that compare Fable 5.1 and some composite benchmarks actually show that it’s worse? How can this be true, and how can ARC-AGI 3 be 99.9%? Is this a path towards AGI?
Let’s review key benchmarks like DeepSWE, ExploitBench, ARC-AGI-3, and more.
#openai #llm #ai
Follow me:
X:
LinkedIn:
TikTok:
Chapters
00:00 Intro
00:16 Signal vs Noise
01:30 ARC-AGI-3
03:56 Focus
04:49 FrontierMath
05:34 DeepSWE
06:12 Sponsor: Zo
07:20 Pareto Frontier
08:32 Macro Trend
09:52 Computer Use
10:35 Good model?
コメント