Executive Summary
OpenAI launched GPT-5.6 Sol and Terra in a limited preview, but the release was overshadowed by Claude Mythos 5 outperforming Sol on cybersecurity benchmarks—prompting OpenAI to reframe its results around "performance-efficiency" rather than raw capability. Simultaneously, the open-source AI ecosystem accelerated dramatically: a new 9B-to-397B MIT-licensed agentic-coding model family dropped, GLM 5.2 matched Claude Code in independent testing, and DeepSeek V4 was successfully pruned to run on a single desktop DGX Spark unit. Hardware constraints—particularly memory bandwidth on local devices—emerged as the defining bottleneck, with dual-Spark interconnect setups promising to double throughput. Regulatory headwinds loom as GPT-5.6 enters a government review period, fueling a surge in local/open-source AI adoption sentiment.
Key Events
- OpenAI releases GPT-5.6 Sol and Terra in limited preview, but benchmark charts show Claude Mythos 5 leading at ~80% vs Sol's ~73%; OpenAI repositioned the axis as "performance-efficiency," drawing criticism. → link
- New open-source agentic-coding model family released spanning 9B to 397B parameters under MIT license; the 35B MoE variant fits a single RTX 3090. → link
- DeepSeek V4 (180B via REAP pruning) runs on a single DGX Spark, with detailed benchmarks showing ~11-12 tok/s sustained generation across both ds4 and vLLM engines, limited by memory bandwidth. → link
- GPT-5.6 enters a government review period, with restricted access during evaluation before potential public release—raising concerns about regulatory capture and frontier model restrictions. → link
- Apple M5 Max 128GB MacBook Pro price jumps 29% from $5,399 to $6,999, driving urgency for local-AI hardware procurement. → link
- Hermes Agents updates: Kanban improvements with typed block reasons and recurrence counters, plus a new subagent delegation tool and community-built docs search. → link
- GLM 5.2 reportedly matches Claude Opus 4.8 in OpenClaude coding benchmarks, signaling competitive open-source alternatives. → link
Analysis
Pattern: Efficiency reframing as capability gaps widen. OpenAI's pivot from capability benchmarks to "performance-efficiency" mirrors a broader industry dynamic—when frontier labs can't claim the top capability spot, they reframe the metric. This happened with Sol's cybersecurity scores and will likely recur as open-source models close gaps.
Pattern: Local AI is hitting a bandwidth wall, not a compute wall. Multiple independent tests confirm that running 180B+ models on desktop hardware is feasible, but memory bandwidth (273 GB/s on DGX Spark) becomes the hard ceiling. The fix is dual-unit interconnects, making bandwidth the new bottleneck to watch.
Pattern: Regulatory pressure accelerating open-source adoption. The GPT-5.6 review period and parameter-based restriction proposals are driving developers toward local, open-source alternatives. Predictions of "open-source AI mass adoption within 18 months" are increasingly grounded in hardware availability and model quality rather than ideology.
What to watch next: Dual DGX Spark benchmarks at depth; the full open-source 9B-397B model family's real-world performance; whether GPT-5.6's review period expands or contracts access; and REAP pruning adoption for making frontier-scale models run locally.
Tweet Feed
AI Model Releases & Benchmarks
@sudoingX · 2026-06-26T18:58
the most funded lab on earth, racing to second place on its own slide. openai dropped "our most capable cybersecurity model," then published the chart themselves: the top line is Claude Mythos 5 at ~80%, their flagship Sol caps at ~73%... they renamed the axis "performance-efficiency" and sold "we reach our lower number in fewer tokens" as the win. → tweet
@gdb · 2026-06-26T17:13
GPT-5.6 Sol preview — it's a good model: https://t.co/UihzcpfR22 → tweet
@jezell · 2026-06-26T17:48
RT @OpenAI: Introducing a limited preview of GPT-5.6 Sol, our next generation frontier model, as well as GPT-5.6 Terra, a balanced model fo… → tweet
@kunchenguid · 2026-06-26T17:55
this is very exciting - great to see OpenAI not only pursuing intelligence but also pushing cost efficiency at the same time so the intelligence becomes more accessible and economically viable for real stuff → tweet
@jezell · 2026-06-26T17:54
RT @Hesamation: HOLY SHIT… GPT-5.6 Sol scores just as strong as Mythos Preview at 1/3 of output tokens, on cybersecurity and it's designed… → tweet
@sudoingX · 2026-06-26T18:35
three years of 4o, o1, o3, 4.5 training us to accept the most confusing naming in software, and the grand fix is Sol, Terra, Luna. we went from a math exam to a horoscope. → tweet
@thdxr · 2026-06-26T01:58
the news about 5.6 is being phrased in a certain way but i think it's going to be a thing where it's in a review period for a bit during which access is restricted (because it's unreviewed) and then it'll be made public like any other model → tweet
@juliarturc · 2026-06-26T05:33
Good thing they're banning GPT 5.6, it already caused an earthquake yesterday → tweet
Local AI & Hardware Benchmarks
@sudoingX · 2026-06-26T17:00
i ran the boring tests so you don't have to lose 6 hours figuring out how to run DeepSeek V4 on a DGX Spark. here's the whole thing start to finish... ds4: dead flat ~12 tok/s... vLLM: ~24 on an empty context, drops to ~11... on real work with real context loaded, both engines generate at the same speed, ~11-12, because a 180B model is bottlenecked by memory bandwidth. → tweet
@sudoingX · 2026-06-26T13:11
the deepseek v4 on a dgx spark story got more honest overnight... i put it to actual work, wired it into a hermes agent autonomous running real tasks with tools, and the sustained speed settled at 7 to 8... the gap between a benchmark and a working agent is the most useful number i can hand you. → tweet
@sudoingX · 2026-06-26T17:22
spent the week benchmarking DeepSeek V4 on a single spark and kept hitting memory bandwidth wall. a 180B model reads weights faster than one spark feeds them, so gen caps ~11 tok/s no matter the engine. @MiaAI_lab is pointing right at the fix: two sparks over the interconnect, 40 tok/s, 1M context. → tweet
@sudoingX · 2026-06-26T17:51
i spent a week proving a single dgx spark caps out on memory bandwidth. the cure is a second one wired over connectx, i have neither the second spark nor the cable, just a thirst only nvidia can quench. → tweet
@sudoingX · 2026-06-26T14:26
there's a 180 billion parameter model running on the dgx spark sitting on my desk right now... the frontier didn't stay locked inside three companies' datacenters, it quietly got small enough to sit on a desk. we're going to look back at renting intelligence by the token the way we look back at renting time on a mainframe. → tweet
@sudoingX · 2026-06-26T17:59
funny thing. i spent the whole week documenting the dgx spark's limits... and my dms are filling up with people telling me they just bought one. turns out honest numbers move hardware better than any spec sheet. → tweet
@sudoingX · 2026-06-26T18:22
you will see dgx 128GB box and think "nice, i can finally run a 70B." nah. that 128GB is a MoE monster machine anon. and the entire wave of big MoE models, deepseek, glm, qwen, the new ones landing weekly, those are exactly what this box was built for. → tweet
@alexocheema · 2026-06-25T21:54
The 128GB M5 Max MacBook Pro went from $5399 to $6999 (+29% more expensive!) A customer of ours bought 42 of these just before the price increase - they saved $67,200. If you want to run local AI, you need to secure hardware ASAP. → tweet
@TheAhmadOsman · 2026-06-26T05:48
I am not kidding, now is the time more than ever to hunt an RTX 3090 and learn how to run Qwen 3.5 27B → tweet
Open Source AI
@sudoingX · 2026-06-26T18:27
open source is flying right now. a full agentic-coding family, 9B to 397B, MIT licensed, dropped like it's a tuesday. the one i actually want to run is the 35B MoE, it fits a single 3090 and walks onto a dgx spark. → tweet
@TheAhmadOsman · 2026-06-26T04:27
Prediction: Opensource AI mass adoption within less than 18 months → tweet
@TheAhmadOsman · 2026-06-26T00:48
Thanks to GLM 5.2, I know for a fact that enterprises are moving off the cloud, acquiring compute, and working on having post-trained models for their own use cases. It's checkmate for Opensource AI, they just don't know it yet. → tweet
@TheAhmadOsman · 2026-06-26T00:02
Opensource AI needs Local AI to survive. Once people get their heads around that we will all be better off. → tweet
@TheAhmadOsman · 2026-06-26T02:11
We have entered the Local and Opensource AI rocket acceleration phase. Next week is gonna be wild. → tweet
@victormustar · 2026-06-26T12:52
nature is healing... it might already be over for closed AI → tweet
@sudoingX · 2026-06-26T08:56
opensource must win. nous research must win. → tweet
@victormustar · 2026-06-26T11:18
300tok/s on mobile is insane... open source must win ✊ → tweet
Regulatory & Policy Concerns
@tinygrad · 2026-06-26T17:18
Someday we'll look back at when the government tried to restrict AI with too many parameters like when they tried to restrict encryption with too many bits. Though unlike last time, this time there will be consequences. The world will see China as the clear AI leader. → tweet
@Ex0byt · 2026-06-26T03:02
Suppression achieves its own indictment. Forbidden loses to available. ...Prohibition didn't end drinking, it bore Capone. Throttled frontiers won't abolish the open source, it will hand it to someone else (Chynaa), leaving the best models locked in a drawer while good-enough ones ship everywhere else. → tweet
@Ex0byt · 2026-06-25T21:27
OpenAI's next model ships in staggered preview, with the government playing clearing house - "voluntary", for now. my open source services will be in even higher demand. → tweet
@thdxr · 2026-06-26T13:40
the self-own that's happening to the ai industry right now is a great reminder of why human brains are just as important as ever. all the compute in the world and they couldn't foresee this basic situation. ai not gonna save you from having to be competent → tweet
@thdxr · 2026-06-26T13:40
you know how sometimes you say stuff but don't expect anyone to actually take up the offer... this is what they were doing with the "we should be regulated" bit → tweet
@TheAhmadOsman · 2026-06-26T03:24
Anthropic losing is as sweet as Opensource AI winning and I know that sounds weird but those guys are evil → tweet
@TheAhmadOsman · 2026-06-25T21:04
Prediction: Karpathy leaves Anthropic within less than 6 months from now → tweet
@TheAhmadOsman · 2026-06-25T19:14
Continual Learning will run locally. That's why the big labs aren't talking about it. Not your weights, not your model, LITERALLY → tweet
@TheAhmadOsman · 2026-06-26T08:09
Cry me a river you trained your models on humanity's knowledge → tweet
Developer Tools & Agent Frameworks
@Teknium · 2026-06-26T17:28
Hermes Agent's Kanban got a bit better today! Added typed block reasons and a recurrence counter to Kanban tasks so that repeated same-cause blocks escalate to a human instead of looping forever, while dependency blocks auto-resume when their parent completes. → tweet
@Teknium · 2026-06-26T17:18
Check out Tonbi's latest video to learn about Hermes Agents' subagent delegation tool! → tweet
@thdxr · 2026-06-26T14:03
mcp was clearly inspired by looking at lsp and all the progress in the spec has been slowly realizing it should be nothing like lsp → tweet
@thdxr · 2026-06-26T16:40
lot of people saying claude tag is just a slack bot and they're exaggerating but i think they're right on this one, it completely changes how your whole team works and it's so fun! everyone's addicted → tweet
@kunchenguid · 2026-06-26T05:45
i have stopped using fast mode with gpt 5.5. with firstmate handling context switch, my average number of parallel sessions went from 3-5 to now 10-20. making requests faster (and more expensive) doesn't help at all when token quota is the clear bottleneck → tweet
@ASalvadorini · 2026-06-26T07:09
In case anyone wants to see how to use Minimal, I updated my Flutter Architecture Component repo to use Minimal 3.0.0, and surprise surprise while at it I introduced a new release agent 😅 → tweet
@gdb · 2026-06-25T23:00
digital ocean for running a codex remote session: https://t.co/IdCkX3QgOf → tweet
Research & Technical
@Ex0byt · 2026-06-26T17:54
10k/hrs of per-frame video annotation. now that's what we call a dataset. Nice work Reka! and thanks for open sourcing this data gold mine. → tweet
@Ex0byt · 2026-06-26T14:44
Interesting... tree-causal attention mask inside a single parallel draft pass. is it useful beyond easy query/single pass inference? JetSpec: https://t.co/pSYGvrM6dH → tweet
@Ex0byt · 2026-06-25T20:42
will play around with this and see if we can produce the best multi-pass NVFP4 and a PRISM-REAP for ya'll. → tweet
@Ex0byt · 2026-06-25T22:03
GOLD MODEL! And fits with my JIT Expert Prefetching → tweet
@jsuarez · 2026-06-25T19:43
Reinforcement learning research with Joseph Suarez https://t.co/ccCmKqjaJ3 → tweet
@thdxr · 2026-06-26T02:25
the reason we excluded frontier models from our data page is they are artificially low in usage and didn't want people concluding this... most people using them use them directly, they don't buy them through us. we wouldn't be surprised if they're 90%+ of total token spend → tweet
@victormustar · 2026-06-26T16:55
trained my first Krea 2 LoRA - super impressed by it 😙 → tweet
Open Source Tools & Projects
@jezell · 2026-06-25T19:25
RT @ClickHouseDB: Reliable backups are foundational to running Postgres in production. That's why we've open-sourced WAL-RUS, a Rust imple… → tweet
@cooltechtipz · 2026-06-26T17:16
TPU vs GPU https://t.co/kfKBymqhs8 → tweet
@cooltechtipz · 2026-06-26T08:32
How to build a local AI https://t.co/kabBZp4sQ5 → tweet
@cooltechtipz · 2026-06-26T04:15
How open frameworks and open data ecosystems work together. https://t.co/kChRfoBoO2 → tweet
@RydMike · 2026-06-25T20:36
Check out this awesome 🎉 chart 📈 pkg for #FlutterDev... maybe we have a new charting champ in Flutter. → tweet
@LinusEkenstam · 2026-06-26T09:52
RIP Adobe & Claude Design .md. Designers are reborn → tweet
Industry Commentary
@thdxr · 2026-06-26T13:40
you know how sometimes you say stuff but don't expect anyone to actually take up the offer... this is what they were doing with the "we should be regulated" bit → tweet
@sudoingX · 2026-06-26T07:23
hot take: once someone joins closed frontier lab, i start weighting their incentives as heavily as their opinions. that's true for everyone. → tweet
@victormustar · 2026-06-26T12:12
a bit more than just another AI company https://t.co/pOPAaAsZ1n → tweet
@ASalvadorini · 2026-06-26T06:21
With my Gemini Pro subscription and my use, I can use Opus 4.6 for 1.5 hours a day for 3 days a week, until the quota resets. The Gemini models instead are generous enough to never hit the limits (so far). → tweet
@LinusEkenstam · 2026-06-26T09:17
This guy has single handedly made sure AI is so powerful and widely used that even my grandma now uses Claude code → tweet