Executive Summary
The last 24 hours saw a major shakeup in the AI landscape with the widespread rollout of GPT-5.6 "Sol," which sets a new state-of-the-art on ARC-AGI-3 and deeply integrates agentic capabilities into Microsoft 365 and ChatGPT Work. Simultaneously, Grok 4.5 entered the coding model arena with highly aggressive pricing, igniting a price war for developer-focused models. On the local AI front, developers are heavily optimizing model quantization (such as Qwen 3.6 and MLX-based models) and leveraging high-unified-memory hardware like the Framework Desktop 128GB and NVIDIA's DGX Spark to run frontier-level models entirely offline. Open-source contributions remain robust, highlighted by MiniMax's commitment to open-weight frontier models and new releases from AntGroup and the open-source community.
Key Events
- OpenAI's GPT-5.6 "Sol" becomes the preferred model for Microsoft 365 Copilot and sets a new SOTA on ARC-AGI-3 (7.8%). → tweet link → tweet link
- Grok 4.5 is released as a coding agent with disruptive economics ($2/$6 per million tokens), taking market share from Claude Opus and sparking a price war. → tweet link
- ChatGPT Work and the new ChatGPT desktop app unify Codex under a cloud harness, bringing agentic computer use to consumer scale. → tweet link → tweet link
- MiniMax CEO announces personal transition to zero compensation, pledging 1% of shares to a dedicated fund for open-source communities and the broader AI ecosystem. → tweet link
- A developer successfully ports 2 million lines of LibreOffice code to Dart using OpenAI Codex and GPT-5.6. → tweet link
- Framework Desktop 128GB proves highly capable as an AMD-based local inference box, running a 397B parameter MoE model at ~18 tokens/sec using open Vulkan drivers. → tweet link
Analysis
The AI ecosystem is rapidly shifting from conversational chat to autonomous agentic workflows, with Codex and Hermes Agent leading the charge in developer tooling. The release of Grok 4.5 at a fraction of the cost of Anthropic's Opus models indicates an aggressive price war, potentially pressuring Anthropic to adjust its compute efficiency and modality strategies. Meanwhile, the local AI movement is gaining unprecedented traction; developers are proving that dense 27B-35B models (like Qwen 3.6) running on 24GB GPUs or high-memory AMD/NVIDIA systems can deliver viable daily-driver performance without API dependencies. Watch for further consolidation in AI agent harnesses, deeper integrations of local models into standard IDEs, and a potential pivot from closed-labs as open-source catches up in benchmark performance.
Tweet Feed
AI Model Releases & Benchmarks
@sama · 2026-07-10T14:18
GPT-5.6 is now the preferred model in Microsoft 365 Copilot
https://t.co/r2B3EJuV1A → tweet link
@sama · 2026-07-10T04:51
RT @arcprize: GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8%
Sol is the first verified frontier model to ever beat an ARC-AGI-3 game
It i… → tweet link
@sudoingX · 2026-07-09T20:00
what a move by elon. everyone else spent two years scraping github for code, xai went and trained the model with cursor, directly inside the place where real engineering happens. distribution and training data in one deal.
and the economics are the actual announcement. $2/$6 per million tokens, ~4x fewer output tokens than opus on the same swe-bench tasks, 80 tok/s. the intelligence board is honestly mixed, it takes terminal-bench and deepswe off opus 4.8, drops swe-bench pro and multilingual. elon himself called it opus 4.7 level. that's the right read.
but here's the thing, if opus level intelligence gets this cheap, the calculus changes. grok 4.5 takes the everyday agent work and my claude code subscription quietly turns into a fable subscription, the button i press when the task is actually hard.
the coding model war just became a price war. price wars are won by users. → tweet link
@RayFernando1337 · 2026-07-10T09:49
RT @Lbtechreal: Initial impressions of Grok 4.5 as a coding agent: really, really good.
We pointed firstmate + Grok 4.5 at MatildaOS (priv… → tweet link
@kunchenguid · 2026-07-09T22:55
right now i have fable still in the subscription with limits just reset; i have gpt 5.6 just added to the mix; and i have grok 4.5 also working super well
after using all of these for a while, i suspect something really interesting is about to happen
tl;dr - i think anthropic is in serious trouble
- portfolio competitiveness
opus and sonnet (who still remembers they have haiku?) are basically not even worth using any more. both gpt 5.6 and grok 4.5 are just outright better
fable is the only real competitive option from anthropic now, but it's not even going to stay in the subscription for a while. it's also super slow and expensive
not having enough focus on efficiency and being obsessed with intelligence upper bound might have just led anthropic down a bad road
- compute constraint
an interesting side effect of grok becoming so good is that its demand might explode, especially with cursor now heavily doubling down on grok
if grok suddenly gets a lot of usage, it will need compute. and xAI may not want to rent compute to anthropic any more
if anthropic can't secure more compute, they will be in trouble because there's just no way to grow the business
- modalities
both openai and xai started investing in multi-modality since the beginning, while anthropic tunnel-visioned into text and code
now i find myself increasingly rely on openai's image generation and grok's video generation. if i can only keep one subscription i would have no choice but to drop anthropic because i do need the other modalities
overall at this rate, what's going to happen is that the vast majority of agentic work will be done by non-anthropic models, and only very occasionally other models will escalate hard problems to fable as an advisor
my advice for consumers remains the same as what I've been saying for a while - get ready for a multi-agent, multi-vendor world. use tools that allow you to freely switch between models, and build a setup where you can use the right model for the right task → tweet link
@sudoingX · 2026-07-10T06:07
meta just shipped its first paid closed source model. the company that carried open weights for the entire west closed the gates the same moment its model started topping benchmarks.
i told you openness was a strategy for second place. i didn't expect meta to prove it this fast.
look at the full board now. america export controls its frontier. china drafts rules to keep its models home. meta goes closed the moment it's ahead. every player, the same move, the same month.
the era of assuming a better open model always arrives next quarter is ending in real time, in public, on a schedule you can watch.
what's on your drive is yours. what's behind an api is a promise. and promises are getting repriced weekly now. download accordingly. → tweet link
Developer Tools & Agents
@gdb · 2026-07-10T17:09
ChatGPT Work brings agents to consumer scale.
It’s both a step up in usability (you can just do things from your phone, no laptop required) and access, for both personal and professional life.
Very excited to see how people put it to use and how it can empower them! → tweet link
@nummanali · 2026-07-09T22:21
The ChatGPT mobile app has finally converged on Codex
Now “Work” is the mode that uses the Codex harness in the cloud and the Codex navigation has been replaced with “Remote”
Overall great improvement and I imagine a lot more to come
Super happy with the updates! https://t.co/yrUQeg0wmf → tweet link
@jxnlco · 2026-07-10T17:07
RT @JHX73307308: codex's internal browser is cracked, being able to carrying over login status and stuff from Chrome is a game changer, not… → tweet link
@jezell · 2026-07-10T16:23
2M lines of code so far in LibreOffice to dart port. Codex is now running on 5.6. /goal just keeps going. https://t.co/zxwO0oqah7 → tweet link
@thdxr · 2026-07-10T17:41
if you want to help beta test OpenCode 2.0 https://t.co/XXFixRHgSw
- data in separate db which we might wipe
- stuff will be broken
- use /report to send us issues
- v1 plugins won't work, v2 api not final
there is a built in skill that you can ask for basically anything → tweet link
@RayFernando1337 · 2026-07-10T16:55
RT @shadcn: Introducing shadcn/typeset.
You know how you render markdown and get back plain, unstyled HTML? Headings, paragraphs, lists,… → tweet link
@levelsio · 2026-07-10T18:07
So I am trying to hit a 500 calorie deficit every day + hit my protein goal of about 150g (2g per kg bodyweight), and I used Claude chat for that, but after a week it starts losing track so I asked it to export its data as CSV and copy pasted that into Claude Code on the VPS
And asked it to build a little calorie tracker called 🥩Caltrack with a Telegram bot too, so I can log whatever I eat or drink in there either via Telegram or if I want more granular via Termius SSH on the VPS
I like to do this mix of Claude Code and a dashboard/chatbot because you can ask more deep questions to it like "how am I doing", "what to improve" etc. it's just much smarter and "able" on the server than the Claude chat app by itself → tweet link
@sudoingX · 2026-07-10T05:49
tmux is the most non-negotiable piece of my entire stack. not the gpu, not the models. tmux. let me show you why anon.
this is my machine right now, 14 named sessions, every one an agent or a service with its own lane.
-
every agent lives in its own session with a real name. reviewer, builder, server, research, content. names are what make delegation real, "check the reviewer pane" is an instruction a human or another agent can follow, and when something breaks at 2am i know exactly which door to knock on.
-
the big agents spawn their own subagents inside those lanes. the reviewer audits a task, writes the brief, hands it to the builder session, then sits there waiting to tear the PR apart. builders fan out their own workers when the task is wide. i read conclusions, not diffs, and the tree of agents underneath stays their problem.
-
one pane serves a 27b model on localhost, the next pane runs the agent that talks to it. the model and its consumers are neighbors in the same house, no cloud between them, and when i want to swap the model i restart one pane and nobody else notices.
-
grok sits in its own pane with live x search, that's the research desk, it pulls timelines and receipts while everything else builds. claude runs planning and the files. cursor agents write the code. three vendors, one terminal, zero browser tabs, and they read each other's work through the panes.
-
no ide anywhere. every serious agent is a cli now, and the terminal quietly became the better ide while nobody was looking. syntax highlighting doesn't ship features, agents do.
-
everything survives me. close the laptop, kill the connection, reboot the router, the fleet keeps its state. attach tomorrow morning and every conversation sits exactly where i left it mid-thought. tmux is the persistence layer of the whole company.
this is not a dev setup. it's an office. every session is an employee that never clocks out. if you run agents and you don't run tmux, you're running them with one hand. → tweet link
Local AI & Hardware
@sudoingX · 2026-07-09T19:47
one month with the framework desktop 128gb as my amd inference box and i keep waiting for the catch. there isn't one. 128gb of unified memory in a 4.5 liter box.
strix halo, ryzen ai max plus 395. framework sent it, amd seeded the program, and my only job was honest numbers wherever they land. so here are the numbers. a 397b moe runs at around 18 tok/s on this thing. a 35b holds a four million token context window. and the open community vulkan drivers beat amd's own rocm stack on my bench, which is the most linux thing imaginable.
three and a half weeks of daily serving now, model swapping, long agent sessions, zero crashes, zero driver tantrums, fans i genuinely cannot hear. it sits next to the nvidia box and eats whatever gguf i throw at it.
my 24gb card is still the daily driver, but a big memory tier at this price did not exist a year ago. amd showed up, and this quiet little box is the proof. → tweet link
@sudoingX · 2026-07-10T06:37
i took qwen 3.6 27b dense, the king of the 24gb tier, for a full bench spin on my rtx 5090 laptop, head to head against a used $900 rtx 3090. some results i expected. one genuinely surprised me.
the setup: q4_k_m, 16.8gb file, llama.cpp with flash attention, q4 kv cache, same flags on both cards, single stream, all measured this week on my own hardware.
what i found:
35.3 tok/s generation fresh, 1,509 tok/s prompt processing on the laptop.
the used $900 3090 beats the flagship laptop at generation, 40.1 against 35.3. generation is memory bandwidth, and the old card simply has more of it, 936 vs 896 GB/s.
the laptop wins prompt processing by 16 percent. blackwell compute is real, it just isn't what generation is hungry for.
filled context is the silent killer. 36 tok/s fresh glides to 18.8 by 128k deep and 12.9 at the full 262k. minus 64 percent. nobody benchmarks this and every agent user feels it.
128k context costs 18.7gb of vram, the sweet spot for agent work. the wall is 376k. and the footprint is identical on a 3090, 4090 and 5090 within 35 MiB.
full picture in the chart, every flag included, everything reproducible.
next up: the depth curve on its own, the measurement almost nobody publishes. stay close. → tweet link
@ivanfioravanti · 2026-07-10T15:22
MLX: playing with Pi, Hy3, oMLX, M3 Ultra 512GB and liteLLM proxy to reach my AI lab from anywhere. 💪
Great model, perfect tool calling and instructions following. Here runs without any specific frontend skills, the goal was just evaluating it in a coding agent environment.
I’ve used the 512 branch (8bit dynamic quantization) of Hy3-Alis-MLX-Dynamic by @Alisvolatprop12 (great job!) It runs at around 15 toks/sec as average on a single M3 Ultra. I'll try distributed soon. PR is still open so performance improvement could come in future. Model used is here on HF: https://t.co/ldpT0EApkV
Thanks @TencentHunyuan for this great model 🙏
Note: initially I was getting many errors on Pi side for request terminated. I investigated and it was a timeout issue on my liteLLM configuration. Local models need more time.
I added this to LiteLLM config.yaml
router_settings: timeout: 600 → tweet link
@Ex0byt · 2026-07-10T17:18
I've opened the private vault. Enjoy!
Qwen3.6-35B-A3B-PRISM-MLX-NVFP4 for MLX running ~90 tok/s (23GB). Our PRISM-Pro Dynamic quant is the best quant out there on accuracy/speed, tool calling, and agent use.
Now Available on HuggingFace.
- (Collection): https://t.co/y7R1SVmnkG
- (Direct link): https://t.co/aSKXKJyDG7 → tweet link
@ivanfioravanti · 2026-07-10T15:12
RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU.
Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 1… → tweet link
@sudoingX · 2026-07-10T07:36
if you've been hesitating on local ai, waiting for some future where it gets good enough, this is the post to hear. it's already good enough, and it has been for months.
one 24gb discrete gpu is the entire entry ticket. a used 3090, a 4090, any card in the tier. load qwen 3.6 27b dense at q4_k_m and you'll feel it inside the first hour. a real model, on your machine, thinking with you, and nobody sitting between you and it. your prompts stay in the room. your thinking stays yours.
that model is the king of the tier, my whole timeline census just voted the same way, and the numbers are on my last chart. the setup takes one evening. the hesitation has already cost you more than the card will.
bookmark this for the day you're ready. it gets cheaper to start every month, and it never gets less yours. → tweet link
Open Source & Ecosystem
@SkylerMiao7 · 2026-07-10T03:32
A letter from our CEO today:
Markets will fluctuate, and external noise will come and go, but our direction remains unchanged.
Being at the forefront of this industry, we have a clear understanding of the pace of technological evolution and the long-term value we are building together.
Starting today, and until the day we reach AGI, I will no longer receive any compensation from the company.
Over the next four years, I will allocate shares equivalent to 4% of the company’s total equity from my personal holdings to reward team members who choose to build this journey with us for the long term and create value together. In addition, I will allocate 1% of my shares to establish a dedicated fund supporting the continued growth of open-source communities and the broader AI ecosystem.
I will devote all my time, energy, and resources to this mission.
This is my long-term commitment as a founder — to our company, our team, and the future we are building together.
We will keep going until we get there.
Intelligence with Everyone.
IO CEO & Founder, MiniMax → tweet link
@sudoingX · 2026-07-10T08:52
been going down a rabbit hole on @AntGroup's open models, the ling and ring families. trillion param flagships at the top, but the interesting part is the small end. ring mini is 16b, ring flash and ling flash are ~100b moe with tiny active params, all MIT, all on hugging face.
so, local ai folks, has anyone actually run these? ling 2.6-flash on a single card, ring mini on a 24gb tier, the 1T on a spark or a cluster. what's the real agentic coding experience, not the leaderboard.
and openrouter users, ring-2.6 was quietly near the top for a while. did you feel it or was it just numbers.
drop your honest take. i'm about to put them on my bench and i want to know where to point it first. → tweet link
@cooltechtipz · 2026-07-10T13:39
Something nobody could've imagined 2-3 years back. All ten of the top open-weight models come from China-based labs. → tweet link
@Teknium · 2026-07-10T08:46
RT @RyanLeeMiniMax: At least for MiniMax, we will keep release frontier open weight Model → tweet link
@ivanfioravanti · 2026-07-10T04:07
RT @Prince_Canuma: 🎉 Congrats to @MosiAI_Official + @Open_MOSS on the release of MOSS-Transcribe-Diarize-0.9B — a genuinely impressive end-… → tweet link
Software Development & Backend
@hnasr · 2026-07-10T16:21
Worker pools and connection pools can be mixed and matched into some interesting backend design patterns.
NGINX has a process worker pool, each worker has a dedicated upstream connection pool, Making sharing challenging and increases overhead of connection establishments but reduces contention.
HAProxy uses a threaded worker pool allowing sharing connections between threads with an access to local thread cache, can still run into contention when accessing the shared pool.
I tell a story of how I discovered this concept in chapter 8 in my book, Root Cause, stories and lessons of two decades of backend engineering bugs, and expand on this further on Appendix A. → tweet link
@RealGeneKim · 2026-07-10T14:26
RT @steren: Today we're publicly launching Cloud Run sandboxes.
Here, I start, execute, and stop 1,000 sandboxes in 5s with an average of… → tweet link
@jezell · 2026-07-10T01:22
RT @criccomini: SlateDB is about to have a pluggable WAL. :) Kafka.. wal3.. chorus.. take your pick.
https://t.co/6vk8XFf0iu → tweet link
@jezell · 2026-07-10T18:00
RT @satleri_sentler: NASAがRustでWASM runtime作ってるらしい https://t.co/YmQtzSsjdo → tweet link