Executive Summary
The past 24 hours were dominated by two landmark developments: OpenAI announced that a general-purpose model autonomously solved a prominent open problem in mathematics (disproving a 1946 Erdős conjecture in discrete geometry), and Alibaba's Qwen 3.7 Max dropped with benchmark-leading results surpassing Opus 4.6 Max across multiple evaluations. In the developer tooling space, Hermes Agent received multiple updates including Grok Build integration, while Zed editor launched Terminal Threads for AI-assisted workflows. On the hardware and infrastructure front, Blackwell GPU prices surged, and SpaceX's S-1 revealed a massive $1.25B/month compute deal with Anthropic through 2029.
Key Events
-
OpenAI model achieves first autonomous solution of a major open math problem — a general-purpose internal model disproved a central conjecture in discrete geometry posed by Paul Erdős in 1946, marking the first time AI independently solved a prominent open mathematical problem. → link
-
Qwen 3.7 Max released, beating Opus 4.6 Max on most benchmarks — Alibaba shipped Qwen 3.7 Max just one month after 3.6, with notable scores on Terminal Bench, MCP use, math (44.5 vs Opus 34.5 on apex math), instruction following, and Humanity's Last Exam. Community calls for open-sourcing the weights. → link
-
OpenAI offers $2M in credits to every YC company — Sam Altman announced the program alongside AGI research acceleration and personal AGI initiatives. → link
-
Cohere releases Command A+ as open source (Apache 2.0) — Cohere's best model yet now available with fully open weights. → link
-
SpaceX S-1 reveals xAI Anthropic compute deal at $1.25B/month through May 2029 — Anthropic paying for Colossus compute capacity, with 90-day termination notice, highlighting massive AI infrastructure monetization. → link
-
Blackwell GPU price increase hits tinybox green v2 — tinygrad's supplier raising Blackwell card prices by +$2,200 each, forcing a price hike; last week to order at old price. → link
-
Hermes Agent gets multiple updates — Skill Bundles for batch-loading skills, 20-40% disk space savings in session storage, Grok Build 0.1 early access, and multilingual support. → link
-
Zed v1.3.5 launches Terminal Threads — Run Claude, Amp, Pi, or any terminal-based AI workflow as managed threads inside the editor. → link
-
Flutter 3.44 released — agentic hot reload, hybrid composition, and expanded governance model announced at Google I/O. → link
Analysis
Patterns: The most striking pattern is the accelerating cadence of frontier model releases — Alibaba shipping Qwen 3.6 then 3.7 Max within a single month suggests the gap between Western and Chinese frontier labs is narrowing or closing. Simultaneously, the OpenAI math breakthrough represents a qualitative shift: AI is no longer just matching human benchmarks but generating genuinely new knowledge, a capability Demis Hassabis reinforced with his AGI-by-2029/30 timeline at Google I/O.
Industry infrastructure consolidation: The SpaceX S-1 filing reveals AI compute as a trillion-dollar market segment, with the Anthropic $1.25B/month deal showcasing how compute providers are monetizing excess capacity. Rising Blackwell prices signal sustained GPU demand and supply constraints.
Developer tooling convergence: Multiple AI coding tools are racing to integrate agentic capabilities — Hermes adding skills/grokked workflows, Zed embedding terminal AI threads, Codex compaction improving UX, and OpenClaw connecting to wearables. The recurring frustration with agents being bad at writing tests (@thdxr) highlights a persistent capability gap even as autonomy increases.
Open source vs. closed tension: A strong undercurrent of advocacy for open-source AI persists, with calls for Qwen 3.7 Max weights to be released, frustration at Anthropic's safety restrictions preventing security reviews, and warnings about vendor lock-in. The local LLM ecosystem continues to mature with comprehensive inference engine guides.
What to watch next: Whether Qwen 3.7 Max gets open-sourced; follow-up peer review of the OpenAI Erdős conjecture result; impact of Blackwell price increases on hardware startups; whether agent test-writing capabilities improve in next-gen models.
Tweet Feed
AI Model Releases & Breakthroughs
@sama · 2026-05-20T20:53
a general-purpose model solved a major open problem in mathematics. we'll be saying this a lot over the coming years, but this is a kinda big milestone. i'm very excited for AI to greatly extend our understanding of the world, but still, i have complicated feelings today. → tweet link
@gdb · 2026-05-20T19:32
An OpenAI model has achieved a major breakthrough in mathematics, by disproving a central conjecture in discrete geometry that was first posed by Paul Erdős in 1946. This is the first time AI has autonomously solved a prominent open problem central to a field of mathematics. → tweet link
@gdb · 2026-05-21T07:38
our math result is a milestone in new knowledge generation by AI. very exciting to imagine similar results in other scientific fields. "It's very hard to sleep, man" is a pretty good reaction. → tweet link
@sudoingX · 2026-05-21T18:49
qwen is unreal. they just dropped 3.7 max and it is beating opus 4.6 max on most of the benchmarks they ran. terminal bench, mcp use, math, instruction following, humanity's last exam. and the apex math number, 44.5 against opus 34.5, that is not a small gap. the 35 hours straight on a kernel optimization task with 1000+ tool calls is the part i keep rereading. that is the agent era thing actually happening, not a slide. the speed alibaba is shipping at right now is the whole story, 3.6 was last month, 3.7 max today, nobody else is moving like this. one thing though, please open source this one too. → tweet link
@louszbd · 2026-05-21T09:42
RT @nickfrosst: Command A+ from @cohere is out now :) its our best model yet and its open source apache 2.0 → tweet link
@victormustar · 2026-05-21T08:20
Extremely fast on HuggingChat (served by Cohere) → tweet link
@badlogicgames · 2026-05-21T17:19
RT @nicolaygerold: We made rush actually rush. It is now GPT 5.5 with no reasoning tuned for smaller, bounded coding tasks where higher r… → tweet link
@LinusEkenstam · 2026-05-20T19:42
"We are only a few years away from AGI" — Sir Demis Hassabis. 2029/30 is his current estimates. No hype, just Demis laying out his thoughts on where we are, where we are not, and where we are going. → tweet link
@LinusEkenstam · 2026-05-20T19:25
The Age of AI will be 10x more impactful, 10x faster than the Industrial Revolution. But the future is not written. — Sir @demishassabis → tweet link
AI Agents & Developer Tools
@Teknium · 2026-05-20T19:39
Skill Bundles in Hermes Agent allow you to predefine a set of skills, then force load them all with one slash command. Makes workflows that rely on a set of skills super easy! Get access to them early with
hermes update→ tweet link
@Teknium · 2026-05-21T03:06
Our database and data engineering expert @yoniebans made some major improvements to the way sessions are stored and accessed. This will save something like 20-40% of the disk space used by Hermes Agent to operate, speed up session loading, and overall makes the codebase cleaner, simpler, and better architected! → tweet link
@Teknium · 2026-05-21T01:54
RT @NousResearch: Grok Build 0.1 is now available for early access in Hermes Agent → tweet link
@Teknium · 2026-05-21T04:56
We support a variety of languages in both slash command outputs and the dashboard ^_^ → tweet link
@Teknium · 2026-05-21T02:24
Anthropic's terrible safety situation is making it so that I cannot have Opus review p0 issues in Hermes Agent to review and help fix security issues. This does nothing but give hackers an asymmetric advantage over everyone - they will find jailbreaks, they will find ways around this to exploit systems - and the rest of us are locked out of using AI to protect from them. What a joke → tweet link
@thdxr · 2026-05-20T23:41
the latest flavor of agents being bad at writing tests is instead of testing a module, they do a dummy re-implementation of it and call that. wonder what's making them particularly weak at this → tweet link
@thdxr · 2026-05-21T05:15
had opencode build a way to declaratively apply a schema to a sqlite db in any state - first pass iterated table by table and tried to mutate db to get it to match - i told it to try an ast, diff, apply method - it did that but the public apis were ugly - told it to offer a diff(db, schema) -> ops and apply(db, ops) api. it did most of the work but needed me to make something great → tweet link
@thdxr · 2026-05-21T18:08
RT @xai: You can now use your @grok or X Premium subscription in @opencode. Use the model powering Grok Build for high speed and codebase… → tweet link
@steipete · 2026-05-21T11:08
RT @labibrahman: codex compaction is arguably the single biggest user experience improvement in ai in the last 6ish months and does not get… → tweet link
@RydMike · 2026-05-21T10:23
RT @bridgemindai: New CursorBench results just dropped. Two big takeaways. Composer 2.5 is way better than most people think. 63.2% sc… → tweet link
@badlogicgames · 2026-05-20T23:04
RT @zeddotdev: Terminal Threads are live in Zed v1.3.5! You can now run claude, amp, pi, or any terminal-based workflow as a managed thr… → tweet link
@badlogicgames · 2026-05-20T23:05
ok, maybe it's time to switch to Zed. → tweet link
@steipete · 2026-05-20T19:34
RT @twistartups: Someone connected Open Claw to their Meta Ray-Bans. Shawn looks at a box of lens wipes and tells his agent to add them to… → tweet link
@nummanali · 2026-05-21T18:37
RT @gakonst: Open Sourcing Centaur: Multiplayer, self-hosted, secure agents for Slack. → tweet link
@Ex0byt · 2026-05-20T22:57
This really resonates. What can't agents do today with the right harness? → tweet link
@TrungTPhan · 2026-05-21T18:35
RT @bearlyai: Circle CEO Jeremy Allaire says he used AI to build a "CEO Priortizer". When he receives a request for time, the AI agent scor… → tweet link
Hardware & AI Infrastructure
@tinygrad · 2026-05-21T18:52
Our supplier is raising the price of Blackwell cards +$2200 each. As a result, we'll have to further raise the price of the tinybox green v2 blackwell. This is the last week to get your order in at the old price. → tweet link
@tinygrad · 2026-05-21T16:27
.@tenstorrent when you are ready, we'll get you on MLPerf for $10M. Ground up stack, one 1MB pip install, zero C++20 (pure Python). From how many people I see on X trying to rewrite it, your current software approach isn't working. → tweet link
@TrungTPhan · 2026-05-20T21:43
SpaceX S-1 dropped and its centred on 3 business lines: LAUNCH, CONNECTIVITY, AI: 1GW of compute, 550m active users. [...] Anthropic is paying xAI $1.25 billion per month through May 2029 to use Colossus for compute. [...] TAM is $28.5 trillion. → tweet link
@jezell · 2026-05-21T01:19
RT @deok_filho: Who says Kubernetes can't handle hundreds of millions of sandboxes? → tweet link
@alexocheema · 2026-05-21T00:02
RT @exolabs: We're benchmarking every model, every quant, on every different hardware setup for every price point. All developers, compan… → tweet link
Open Source AI & Local LLMs
@TheAhmadOsman · 2026-05-20T22:08
Sama trying to lock-in as many customers as possible for the next 1-3 years makes way more sense now. Great numbers to show before the IPO → tweet link
@TheAhmadOsman · 2026-05-21T07:38
There are way too many parallels between dictatorships and closed source AI btw. Opensource AI winning is existential. No, I am not being dramatic, consider that you might be really underestimating things please → tweet link
@TheAhmadOsman · 2026-05-21T05:38
My beliefs re: AI - Opensource AI will win - AGI will run locally, not on someone else's server - The real ones are already learning how it works. So: Be early, Buy a GPU, Get ur hands dirty, Learn how it works, You'll thank yourself → tweet link
@TheAhmadOsman · 2026-05-21T04:27
Local AI Is Now Easy With This. Give Codex Cli the article below & tell it: Infer the right Inference Engine from your hardware, Use uv+venv, Pick the right kernels, Tune flags, batching, KVCache, etc, Optimize for your hardware & chosen model. See? SO EASY → tweet link
@TheAhmadOsman · 2026-05-21T01:14
DROP EVERYTHING. The bible for running LLMs locally is now available online to read for free. Covers what to use on Laptop/edge, Mac-first, Single RTX, Multi-GPU, production serving, long-context/MoE. Software: llama.cpp, MLX, ExLlamaV2/V3, vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo → tweet link
@TheAhmadOsman · 2026-05-20T19:50
Inference Engines and why they matter when you're running models locally. High-level overview article that, in my honest opinion, should be read by everyone in tech → tweet link
@TheAhmadOsman · 2026-05-20T21:46
Especially in the AGE of AI - Nice piece of software, is it Opensource? If the answer is no, I usually throw it away in the trash. You cannot build any of your workflows around something that could be rug-pulled from you → tweet link
Software Development & Frameworks
@ASalvadorini · 2026-05-20T19:34
RT @FlutterDev: Flutter 3.44 is here! Agentic hot reload, Hybrid Com… → tweet link
@MengTo · 2026-05-21T14:04
Day 88 vibe coding my AI video editor with SwiftUI. It's insane how Apple provides everything, like removing the background of a video without using a single token. → tweet link
@badlogicgames · 2026-05-20T22:29
for windows terminal to report shift + tab in a sensible way, pi has to call 2 win32 API functions. lazily, i used koffi to do that. which added a 27mb. yesterday on stream, we figured out a demented way to cut that down to 3kb. self-contained shared lib, no CRT dep. → tweet link
@badlogicgames · 2026-05-21T09:22
someone at the foundation doesn't seem to like me. first they stopped sponsorship, now i'm no longer part of the GH org. end of an era i guess :) fine tho, i haven't really contributed anything meaningful directly to the repo. → tweet link
@thdxr · 2026-05-21T14:24
at the stage where most of my day is configuring new SAML apps in google workspace → tweet link
@hnasr · 2026-05-21T14:03
If you found yourself hoping a fix works, know that part of you doesn't understand the problem very well. That part of you wants the bug gone. Resisting the bug creates fear. If you wish to hope, hope for an error so you can understand why what you applied didn't work. This is the art of trial and error. → tweet link
@badlogicgames · 2026-05-21T06:27
yeah, we were sitting at a coffee shop in vienna, and i said: let's write a shitty python web app framework. armin was like "that's a brilliant idea, but i don't know how to program". so i went ahead and taught armin, which explains the python. the rest is history. → tweet link
ML Research & AI Safety
@jsuarez · 2026-05-21T17:23
Reinforcement learning research with Joseph Suarez → tweet link
@jsuarez · 2026-05-20T20:01
Apparently I changed PER from sampling without to sampling with replacement by mistake in Puffer 4. This likely made our implicit curriculum learning weaker than it should have been. Investigating on stream now! → tweet link
@juliarturc · 2026-05-21T18:37
This is why we need to gate-keep science. Two pseudo-intellectuals thinking they discovered something deep, conflating social attention, transformers attention, quantum physics observer (attention). These have nothing in common, other than the ambiguity of English language. Naked ladies on Instagram have nothing to do with a weighted average followed by softmax. But they're both so mind-blown by their discovery. Dunning–Kruger will only get amplified by AI sycophancy. → tweet link
@badlogicgames · 2026-05-21T16:21
RT @michellechen: reverse engineering openai's goblin problem: we took open models and trained them with RL to talk about goblins an exper… → tweet link
Tech Industry & Startups
@sama · 2026-05-20T21:56
three of the things we are most excited about: 1. AGI accelerating research 2. AGI accelerating companies 3. personal AGI accelerating everyone in achieving their goals. today it was great to announce the unit distance result. yesterday it was great to announce that we are offering to invest $2M in openai credits into every YC company. → tweet link
@kunchenguid · 2026-05-21T01:54
as promised, my layoff explainer video is live! did my best to cover the underlying driving forces of tech layoffs, what's going on with Meta, and how we should think about our career → tweet link
@TheAhmadOsman · 2026-05-21T15:46
Ahmad, which company were you talking about? - They no longer train models - Transitioned into a Neocloud that rents compute. "Elon has the mandate" clowns in shambles → tweet link
@louszbd · 2026-05-21T09:41
RT @0xJuliechen: I'm happy to announce that I made it to YC building with @photon_hq residents! → tweet link
@victormustar · 2026-05-21T10:08
RT @Gradio: Something small is coming. 16 days to go. 2 weekends to build. $25k+ in cash prizes. → tweet link
@TheAhmadOsman · 2026-05-21T14:40
OpenClaw was such a weird phase in LLMs. Glad that thing disappeared as fast as it took the spotlight → tweet link
@badlogicgames · 2026-05-21T16:21
RT @adalq: the most Google thing is there is not pricing information on the pricing page → tweet link
@LinusEkenstam · 2026-05-20T21:07
Demis react on stage at Google I/O on the singular moment when him and the Alpha Fold team made that one decision that changed everything and led to the Nobel Prize. → tweet link
@gdb · 2026-05-21T06:02
ChatGPT team ships → tweet link