Executive Summary
This period saw notable activity in developer tooling and local AI deployment. A critical observation emerged around Google’s Gemma 4 12b model exhibiting failure modes in extended agentic tasks—looping into token exhaustion mid-work—while Chinese open-weight models (Qwen) complete identical tasks reliably. The developer community continues shifting toward designing agent loops rather than single-shot prompting, with new tools like "/no-mistakes" (a skill for Claude Code and Codex) and Open Code Review (Alibaba’s AI-powered CLI) reflecting this trend. Hardware benchmarking on the Framework Laptop Max+ 395 demonstrates competitive local LLM performance at roughly half the cost of Apple M5 Max, reinforcing the viability of open-weight models on consumer-grade hardware.
Key Events
- Gemma 4 12b agent loop failure documented in detail — A practitioner ran a 10-minute n-body physics task through an agent harness; the model correctly implemented a symplectic integrator and stable orbital physics, then fell into a 34,000-token loop during testing and could not self-correct. Comparable tasks complete cleanly on Qwen 2.5 27b. → tweet
- "/no-mistakes" agent skill released — Popular agentic engineering tool now invocable as a slash command in Claude Code, Codex, and compatible platforms. Designed to validate agent output after code changes are made. → tweet
- Framework Laptop Max+ 395 benchmarks Qwen3-TTS at ~50% the cost of M5 Max — Detailed comparison across MLX (macOS), GGML Metal, and GGML Vulkan backends shows the Framework laptop delivers comparable TTS inference performance at under half the price, with notes on required GGML kernel optimizations. → tweet
- Community consensus: prompt agents, then design loops over them — Multiple practitioners (steipete, kunchenguid, thdxr) reinforce that single-shot prompting is "so 2025" and the paradigm shift is toward architecting self-prompting agent loops. → tweet
- Nous Research Hermes desktop app gains traction — Teknium (key open-source model contributor) shares positive sentiment about the Hermes agent desktop app as a significant upgrade over prior interfaces. → tweet
Analysis
Agent Architecture Over Prompting. The community is coalescing around loop-based agent design as the dominant paradigm for 2026. Multiple posts from practitioners (steipete, thdxr, jezell) signal that single-response prompting is outdated; the shift is toward building self-referential loops where agents prompt themselves. The "/no-mistakes" tool and Open Code Review are concrete examples of tooling built for this loop-centric workflow—validation and review stages within iterative agent cycles.
Open-Weight Model Reliability Gap. The detailed Gemma 4 12b failure report is the most technically significant observation of the period. A model capable of correctly implementing a symplectic integrator (a non-trivial physics choice) then looping into token exhaustion during a test-writing phase reveals a specific failure mode: models that can plan and execute complex tasks may lack the grounding to recognize task completion and exit. This contrasts sharply with Qwen series models, which practitioners report reliably complete tasks and stop. This reliability gap has material implications for agentic deployments—choosing the wrong base model can mean agents never finish work autonomously.
Local AI Hardware Maturing. The Framework Laptop benchmark demonstrates that consumer-grade hardware (sub-$2,000 laptops) can now serve demanding LLM workloads (Qwen3-TTS inference) at performance parity with high-end workstations (M5 Max). This reinforces the trend of viable local AI without cloud dependency, though the practical ceiling remains around 2 concurrent users for real-time TTS.
Rust for LLM/Agent Tooling. TheAhmadOsman's statement that "Rust is the language for LLMs & Agents" reflects an ongoing shift in the developer community—Rust is increasingly the material of choice for building inference engines, agent frameworks, and performance-sensitive tooling (as evidenced by badlogicgames' Rust MLX C bindings work).
Tweet Feed
AI Model Performance & Observations
@sudoingX · 2026-06-07T17:18
i watched gemma 4 12b build something genuinely impressive today, and then loop itself to death right in front of me. the full run is in the video, sped up but completely uncut, watch it to the end and you will catch the exact moment it stops building and starts looping right in the middle of the work. the task was clean, build a single file gravity simulator, n-body physics, orbits, collisions, running locally on one 3090 through an agent. and for ten minutes it was a joy to watch. it reached for a symplectic integrator on its own, the correct one, the kind that keeps orbits stable instead of spiralling out. real gravity with softening, proper orbital velocities, momentum conserved on collision. the physics was right. the thing actually worked. then on the very last step, writing a few tests to prove its own code, it fell into a loop. not a crash, a loop. it started repeating itself and would not stop. ten more minutes, thirty four thousand tokens into a single answer, the same fragments over and over, until i killed it myself. so it's not that gemma can't code. it did the hard part beautifully. it cannot finish. it cannot hold a long task together without unravelling, and finishing is the entire job in agentic work. here's the part that stings. i run this exact task, same harness, same card, on the chinese open models, qwen especially, and i never see this. they build it, they test it, they stop. every single time. google has the raw capability, you can see it sitting right there in the code, and then the model loops itself to death on a task a 27b from alibaba finishes clean. open weights, apache 2.0, so much to love on paper. i just need it to know when to stop talking. → tweet
@sudoingX · 2026-06-07T17:30
here's just the looping part on its own, real time, no edits. gemma 4 12b had already built the thing, the work was done, and instead of stopping it just kept going. the same fragments over and over, talking to itself while the token counter climbed and climbed. watching a model do this is its own special kind of pain, you can see it's stuck but it has absolutely no idea that it's stuck. genuine question for anyone running gemma 4 locally through an agent, are you hitting this too, or did i just pull a cursed seed and everyone else is fine. i can't tell yet if it's the model or me, so tell me what you're seeing. → tweet
@TheAhmadOsman · 2026-06-08T02:38
I now rank Nemotron 3 Ultra among the top 5 Opensource models out there. Frontier intelligence at home → tweet
Developer Tools & Agentic Workflow
@kunchenguid · 2026-06-08T02:57
/no-mistakes is here! by popular demand i've made the most impactful tool in my agentic engineering setup "no-mistakes" invocable as a skill in Claude Code, Codex et al. just type "/no-mistakes" once your agent has made changes, and watch the magic unfold. details below 👇 → tweet
@kunchenguid · 2026-06-07T21:20
i just had the pleasure of going on @petergyang's latest podcast episode and sharing how i use agents productively. we touched on a lot of practical tips of how i tackle the bottlenecks in planning, review and validation and achieve a constant flow state. check it out! → tweet
@steipete · 2026-06-07T18:58
Here's your monthly reminder that you shouldn't be prompting coding agents anymore. You should be designing loops that prompt your agents. → tweet
@jezell · 2026-06-07T22:17
Was just saying something similar to my team last week. If you are using Codex in request / response mode, that's so 2025. → tweet
@jezell · 2026-06-07T01:00
RT @JoeIngeno: Open Code Review - An AI-powered code review CLI tool https://t.co/eVTIZ88TiU #alibaba #AI → tweet
@gdb · 2026-06-07T19:30
Codex use-cases: "From software engineering and design to data analysis and operations, Codex is becoming an AI teammate instead of just an AI assistant." → tweet
@thdxr · 2026-06-07T21:13
liz says when she hears me voice dictating to opencode i'm nicer than i am when talking to people → tweet
@thdxr · 2026-06-07T17:38
a whole bunch of companies that had good primitives but never figured out DX just got saved by AI. i'm using all these things that were too rough to use before → tweet
Open Source & Local AI
@badlogicgames · 2026-06-07T21:56
Finally had some time to spent with my beautiful @FrameworkPuter, specifically a beefed out Max+ 395 with 128GB of LPDDR5x-8000. Quest: get Qwen3-TTS to work as well as on my new M5 Max. I have 2 implementations with 3 backends for the inference engine. - macOS only MLX (using MLX C bindings in Rust) - macOS only GGML Metal implementation - cross-platform GGML Vulkan implementation. Below table shows results for generating 20 seconds of German with a reference voice on the M5 Max (GGML Metal, MLX) and the Framework (GGML Vulkan). The Framework gives me basically the same performance as the M5 Max at less than half the price. I had to massage GGML a tiny little bit, as its 1D conv kernels for Metal and Vulkan weren't ideal for the Qwen3-TTS decoder architecture. The GGML engine is based on https://t.co/xnmyXKPsA6. The MLX engine is based on https://t.co/xnmyXKPsA6. I forked both of these (plus GGML) and did some performance work on the in the past 2 nights. Shitty forks: [three fork links]. Now my @FrameworkPuter can serve the Pipi bots of all the kids in our hood! Caveat: concurrent TTS is realistically only working for 2 kids at a time. Pibot repo here: [link]. Making of here: [link]. Very happy! → tweet
@badlogicgames · 2026-06-07T22:45
recommended reading. → tweet
@badlogicgames · 2026-06-07T23:04
recommended viewing. → tweet
@TheAhmadOsman · 2026-06-07T22:40
I have been seeing lots of people selling Mac minis recently it honestly feels like a missed opportunity to have gotten them into Local AI if the messaging wasn't with the wrong hardware → tweet
@TheAhmadOsman · 2026-06-07T22:51
Everything I have been recently building has been strictly in Rust. Rust is the language for LLMs & Agents → tweet
@Teknium · 2026-06-07T23:36
RT @trevin: The @NousResearch Hermes desktop app is sooooo good. Huge step up from just using telegram. Biggest quality of life upgrade is… → tweet
@Teknium · 2026-06-07T23:48
RT @IBuzovskyi: HERMES AGENT BUILT A SKILL THAT AUTOMATED AN 8-HOUR WORKFLOW INTO 80 MINUTES. THE AGENT BUILT IT FROM ITS OWN COMPLETED SESSIONS. → tweet
Hardware & Infrastructure
@FrameworkPuter · 2026-06-07T22:34
RT @badlogicgames: Finally had some time to spent with my beautiful @FrameworkPuter, specifically a beefed out Max+ 395 with 128GB of LPDDR5x-8000. → tweet
@FrameworkPuter · 2026-06-08T02:30
RT @hwmedialabllc: Building a Breadboard in a @FrameworkPuter expansion card. Drop a small dev board in the socket, prototype in a slot o… → tweet
@TheAhmadOsman · 2026-06-07T20:14
Trying to hire smart engineers? Add GPUs to your job perks and benefits. I know some will sign on the spot if you tel them they get a GB300 Station → tweet
Miscellaneous / Noise
@levelsio · 2026-06-07T20:18
By far my most autistic project ever 🤓🤓🤓 https://t.co/RRYOCWqRQq https://t.co/gLBcfsMl8v → tweet
@alexocheema · 2026-06-07T21:12
domain acquired. https://t.co/8BV2YBHpuB → tweet
@sama · 2026-06-08T00:25
interesting recursive loop here maybe → tweet
@gdb · 2026-06-07T19:30
interesting → tweet
@Teknium · 2026-06-07T23:53
Thanks! 🙏 → tweet
@TheAhmadOsman · 2026-06-08T01:14
You're welcome, Elon. → tweet
@cooltechtipz · 2026-06-08T00:32
Change your perspective, change your experience. → tweet