Executive Summary
The last 24 hours saw two major model releases: Qwen 3.6 (35B total / 3B active MoE, Apache 2.0, with agentic coding matching models 10× its size) and Anthropic's Claude Opus 4.7 (with significantly improved visual recognition and skill invocation). The open-source community also released Ternary Bonsai, a family of 1.58-bit models enabling an 8B model in just 1.75GB, and SLURPY, a zero-training merge of MiniMax M2.5 + M2.7 that beats both parents on HumanEval. Meanwhile, the Hermes Agent ecosystem from NousResearch continued rapid expansion, and Codex introduced background computer use for parallel agent operations.
Key Events
- Qwen 3.6 released — 35B total parameters, 3B active (MoE), Apache 2.0 license, multimodal, with thinking mode and agentic coding rivaling models 10× its size. Benchmarks show 180 tok/s on RTX 4090 and 91 tok/s on M4 Max via MLX. → link
- Claude Opus 4.7 launched — Available via API and Claude Code (
--model claude-opus-4-7), 1M context, same pricing as 4.6. Users report significantly better skill invocation and pushback on weak ideas, but rate limits hit immediately on Max plans. → link - Ternary Bonsai models introduced — 1.58-bit ternary weights ({-1, 0, +1}) enabling 8B weights in 1.75GB, 4B in 0.86GB, 1.7B in 0.37GB. Runs directly in browser via WebGPU at 290MB. → link
- Codex gets background computer use — Multiple agents can now operate alongside the user on the same machine, enabling parallel agent workflows. → link
- SLURPY model released — A zero-training Convex Hull Gradient SLERP merge of MiniMax M2.5 + M2.7 achieving 89.6% on HumanEval pass@5 (vs M2.5's 85.4%), with every tensor uniquely blended. → link
- Hermes Agent ecosystem expanding — Self-improving agent co-evolving with MiniMax M2.7, new community skills (HuskyLens V2 on Pi 5, browser automation), managed VPS hosting, Web UI, and Hong Kong meetup announced. → link
- Allbirds announced pivot to AI — The shoe company is selling its shoe business and pivoting to buying GPUs to lease to developers, citing the GPU supply gap. → link
Analysis
Patterns: The dominant theme is extreme model efficiency — Qwen 3.6's 3B active params competing with 30B+ models, and Bonsai's 1.58-bit quantization pushing 8B models into sub-2GB territory. This signals the local inference capability gap closing rapidly. Simultaneously, the shift from "AI as pair programmer" to "AI as senior engineer you delegate to" is accelerating, visible in Opus 4.7's pushback behavior and Codex's autonomous computer use.
Escalation: The Hermes Agent ecosystem is exhibiting network-effect compounding — community-built skills, hosting options, and international meetups all within days of launch. The velocity of open-source model releases (Qwen 3.5 → 3.6 in rapid succession) and zero-training merge techniques (SLURPY) are compressing development cycles to days.
What to watch: Qwen 3.6 benchmark sweeps on consumer GPUs in the next 24–48 hours will determine whether it truly displaces Gemma 4 and Qwen 3.5 as the go-to local coding model. Opus 4.7 rate-limiting and Claude Code desktop app stability issues may constrain adoption. The Bonsai 1-bit models deserve close attention — if quality holds at 1.58 bits, the "I only have 8GB VRAM" barrier effectively disappears.
Tweet Feed
AI Model Releases & Research
@sudoingX · 18:15
1.58 bit ternary weights. hear this an 8B model in 1.75 gb. this is the kind of research that makes the "i only have 8gb vram" excuse disappear completely. i've been saying the best model for every gpu tier is the mission. if ternary bonsai holds quality at this compression, the 4gb and 8gb crowd just got a real option. → tweet link
@Ex0byt · 14:19
Let's go Qwen team. It's shaping up to be a busy weekend! This one has agentic coding performance matching models 10× its active size, very strong spatial reasoning density and multimodal capabilities. → tweet link
@Ex0byt · Apr 15 21:56
Meet SLURPY - A Convex Hull Gradient SLERP Designer Child of MiniMax-M2.5 + MiniMax-M2.7. 89.6% HumanEval pass@5 vs M2.5's 85.4% with zero retraining. → tweet link
@Ex0byt · 16:33
PRISM Dynamically Quantized MLX demo model for
model-to-model intelligence transferis up on HuggingFace. Enjoy! (PRISM Convex Hull Gradient SLERP — a.k.a. Slurpy) → tweet link
@victormustar · 09:37
290MB = useful LLM (running directly in your browser via WebGPU) 🤯 Bonsai 1-bit is beyond my comprehension... → tweet link
@Teknium · 13:27
Qwen 27B seems to be Twitter's choice for better local model than Gemma-4 → tweet link
@ollama · 14:07
Qwen 3.6 is here, and open-source! Run it locally with improved agentic coding capabilities. Try it with Claude Code: ollama launch claude --model qwen3.6. Try it with OpenClaw: ollama launch openclaw --model qwen3.6. Run it: ollama run qwen3.6 → tweet link
Claude Opus 4.7
@kunchenguid · 18:04
now - some qualitative observations about Opus 4.7. I shared a new idea with it, and I got GRILLED. "Things I'd push back on or sharpen:" "You're conflating..." "Underspecified pieces that will bite you:" "What I'd actually build first:" "Cut scope, sharpen the constraint, then scale." → tweet link
@kunchenguid · 16:23
ok my first quantifiable data point about Opus 4.7 - it's significantly better than Opus 4.6 at invoking the right skills for the right tasks. still not as good as codex, but a massive change from 4.6. → tweet link
@kunchenguid · 14:59
They haven't enabled opus 4.7 in claude code by default yet. But.. I just tested, you can get it via: claude --model claude-opus-4-7. Have fun! → tweet link
@FinansowyUmysl · 17:44
No i mamy nową wersje Opus 4.7. Benchmarki nie powalają, ale i tak teoretycznie ma być lepiej niż Opus 4.6. Największa poprawa jest w rozpoznawaniu wizualnym. → tweet link
@LinusEkenstam · 15:03
Anthropic does not hold back. The amount of domino bricks in place for this calendar to continue like this is either all powered by Mythos or simply superhuman. → tweet link
@jezell · 17:54
RT @theo: I feel bad dunking on them so much but it's genuinely absurd how bad the new Claude Code desktop app is. → tweet link
@TheAhmadOsman · 11:21
Anthropic is fumbling Claude Code more than Steve Ballmer fumbled Microsoft → tweet link
Local Inference & Hardware Benchmarks
@sudoingX · 17:50
180 tok/s generation on a 4090 with qwen 3.6. if you're on a 4090 and not running this model yet you're leaving performance on the table. 3B active params at that speed is insane for agentic coding. → tweet link
@sudoingX · 17:57
mac users qwen 3.6-35B-A3B hitting 91 tok/s on M4 Max 128gb via mlx in lmstudio, that's solid first numbers. → tweet link
@sudoingX · 17:30
85-100 tok/s on the 3090 with qwen 3.6 already? that's in line with what 3.5 MoE was doing. → tweet link
@tinygrad · 14:27
.@UnslothAI so fast with those @Alibaba_Qwen 3.6 GGUFs! Here's Qwen3.6-35B-A3B on a 7900XTX at 90 tok/s, available right now in tinygrad master. → tweet link
@sudoingX · 17:17
the 5090 mobile has 24gb vram, same class as the 3090. when i benchmark a model on the 5090 and give you the flags and the tok/s, that translates directly to your 3090 at home. 7 models loaded on the 5090 today, hermes agent work i've been cooking for weeks is almost ready to ship. → tweet link
Hermes Agent & NousResearch Ecosystem
@sudoingX · 18:26
this is what happens when model teams and harness teams actually talk to each other. minimax co evolving M2.7 with hermes agent's self improving loop. the compounding is real. if you're still on a generic bloat harness you're missing the network effect that's building here → tweet link
@SkylerMiao7 · 12:09
Hermes is just the start. More agents powered by MiniMax are on the way to MiniMax Agent. → tweet link
@SkylerMiao7 · 12:04
MaxHermes is here. Run a self-improving agent in one click — cloud-hosted, secure, and frictionless. → tweet link
@Teknium · 13:09
Excited to see where our partnership leads us @MiniMax_AI! → tweet link
Codex & Computer Use
@badlogicgames · 17:28
background computer use is some black magic. → tweet link
@nummanali · 17:47
Computer Use was released for the Codex App. It also works with the Codex CLI! → tweet link
@gdb · 05:06
always a real feeling of magic to ask codex to perform a task that requires finding information scattered across slack, google docs, notion, and various internal tools, and it just figures it out → tweet link
@nummanali · 17:52
It's happening, we're shifting the narrative from SWE to Senior SWE. "Treat Claude more like a capable engineer you're delegating to than a pair programmer you're guiding line by line" → tweet link
@TheAhmadOsman · Apr 15 22:48
What am I working on? Condensing everything I do into one place: local AI / LLMs, inference + benchmarking, hardware… → tweet link
Developer Tools & Open Source
@badlogicgames · 17:23
update your anthropic api sdk and enjoy the new display field required for opus 4.7+ to spit out thinking summaries. → tweet link
@jezell · 16:42
RT @lancedb: Adding a column to a 10M row dataset: Lance - 13ms ⚔️ Parquet - 520 seconds. → tweet link
@jezell · 18:44
I told you Iron-Proxy was onto something → tweet link
@steipete · Apr 15 18:27
4 months and thousands of work hours later, we have a great security concept; you can go all yolo, use a sandbox (Docker or OpenShell), there are allow-lists and per-access exec allow/deny prompts. → tweet link
@hnasr · 14:02
Any software you build to solve a problem will have a limitation. The danger is in not knowing what the limitations are before building it. The engineer who realizes this and learns the ins and outs of X will be paid 10 times to fix problems caused by X in production. → tweet link
@MatejKnopp · Apr 15 19:21
Function coloring gets a lot of hate, but looking at it it doesn't seem that Zig's async IO approach will be particularly useful in single threaded (main loop) scenario? → tweet link
@MatejKnopp · Apr 15 22:10
There is something magical about running a full-blown Flutter application on web. But it comes at a cost. Jaspr fills the remaining space, and it is a lot of space to fill :) Been following it for years, so happy about the traction it got! → tweet link
@iamdevloper · Apr 15 22:47
Why is it called vibe coding and not 'prompt and pray' → tweet link
Industry & Business
@LinusEkenstam · 13:43
Allbirds, the shoe company announced its selling its shoe business and pivoting to becoming an AI company. They plan to capitalize on the gap in the GPU market, buy GPUs to lend them to developers who can't find enough from Amazon or Microsoft. → tweet link
@LinusEkenstam · 05:34
High signal that we're moving into the early majority adoption territory. Reese Witherspoon just posted about learning AI. If you thought the speed in AI was crazy, it's time to strap in because we're going into ludicrous speed now. → tweet link
@LinusEkenstam · Apr 15 20:52
Tokenmaxxing by @bcherny — $84-120k in token burn in 2 months. This is the new benchmark for tech workers. → tweet link
@LinusEkenstam · 14:26
Your brain is the last private place you have left. The next company to enter it matters more than any app you'll ever download. → tweet link
@carlvellotti · 14:27
Sabi is not what I expected. no surgery, no implants, no Neuralink vibes. just the best non-invasive hardware and dataset in the category. just announced today. → tweet link
Agent Use Cases
@TheAhmadOsman · Apr 15 21:09 (approx)
Agents are helping me SAVE MY WIFE'S DATA. My wife's old MacBook is in the awkward rescue zone. Codex w/ GPT 5.4 Pro is helping me take care of it — explicit preflight checks, command order, stop conditions, transfer fallbacks. → tweet link
@levelsio · 18:11
So I built this situation monitor dashboard of all my projects and it has this AI insights robot that analyzes my business and told me: "Cloudflare ($2,653/mo) has not been reviewed in >1 year despite powering stable 2.9M remoteok pageviews — negotiate volume discount or move static assets to cheaper CDN." → tweet link
AI Video & Creative
@LinusEkenstam · Apr 15 20:45
AI will do so much for old films, colorization was just the beginning. changing video format, adapting to different screens. This was LTX 2.3 on a RTX5060ti. → tweet link