Executive Summary
AI agent orchestration and the practical realities of agentic workflows dominated the tech conversation, with developers sharing hands-on patterns that favor simple terminal-based setups over complex platforms. Open-source AI saw major momentum: Gemma 4 12B demonstrated strong raw coding capability but suffered from catastrophic looping behavior, while the Hermes agent framework rolled out multiple updates. On the hardware front, Nvidia's abandonment of the affordable 24GB GPU tier is creating a supply squeeze that keeps the used RTX 3090 as the default entry point for local AI. Meanwhile, Codex continues to be pushed as an AI teammate for real engineering work, and multiple voices warned against unchecked "tokenmaxxing" culture at companies.
Key Events
- Gemma 4 12B loops to death in agentic coding tasks — @sudoingX documented the model building an N-body gravity simulator flawlessly, then entering an unrecoverable repetition loop during test writing; Qwen models reportedly finish clean on the same tasks → link
- Hermes Agent ships major updates including a desktop app, font customization, and expense-tracking Telegram integration; Teknium signals a "SOUL Hub" and teases a "way too much stacked" release week → link
- Codex being used at scale for real engineering — @jezell is porting LibreOffice via /goal at ~1 billion tokens/day; @gdb says capability overhang is large and the bottleneck is missing context or skills, not model limits → link
- RTX 3090 emerges as the working-class GPU for local AI — @sudoingX argues Nvidia's exit from the affordable 24GB tier makes the 3090 hold value indefinitely while Mac Minis depreciate rapidly → link
- "Tokenmaxxing" criticized as corporate waste — @kunchenguid equates unlimited token budgets to hiring infinite contractors; advocates agent-agnostic HA workflows over single-vendor dependence → link
Analysis
The conversation reveals a maturation shift in how developers relate to AI tools. The initial excitement about "vibe coding" is giving way to pragmatic orchestration patterns: @sudoingX's tmux-pane agent delegation, @steipete's "design loops that prompt agents," and @kunchenguid's agent-agnostic failover all point to a developer class that treats AI as schedulable compute, not a chat partner. The Gemma 4 12B looping incident underscores that open-weight model quality isn't just about capability — completion discipline and stopping criteria are now first-class evaluation metrics, and Chinese open models (Qwen) appear to have an edge here. Hardware remains a bottleneck: Nvidia's pricing gap between consumer and professional tiers has created a secondary market anomaly where 3090s appreciate. Watch for: (1) whether Google patches Gemma's looping behavior quickly, (2) Hermes/SOUL ecosystem maturation, and (3) whether the "simple orchestration" pattern becomes the dominant paradigm over commercial agent platforms.
Tweet Feed
AI Agents & Agentic Workflows
@steipete · 2026-06-07T18:58
Here's your monthly reminder that you shouldn't be prompting coding agents anymore. You should be designing loops that prompt your agents. → tweet
@gdb · 2026-06-07T19:30
Codex use-cases: "From software engineering and design to data analysis and operations, Codex is becoming an AI teammate instead of just an AI assistant." → tweet
@gdb · 2026-06-07T01:48
Whenever I don't use codex for a task, I ask myself why and usually realize that there's some missing context, I needed to write a skill, or I just didn't think to use it. Rarely is it because the task is outside of the capabilities of the model. Overhang right now feels large. → tweet
@sudoingX · 2026-06-07T15:18
if you're just getting into agentic workflows, the orchestration part is way simpler than the people selling you platforms and courses want you to believe. […] you open an agent in every tmux pane you want working for you. you give each one sandboxed skip-permission so it can actually move […] then you spin up one more agent as the main one and you tell it plainly, these other panes are all working with us, your job is to look at them and delegate. that's the whole orchestration. no framework. no dashboard. no per-seat subscription. → tweet
@kunchenguid · 2026-06-07T21:20
i just had the pleasure of going on @petergyang's latest podcast episode and sharing how i use agents productively. we touched on a lot of practical tips of how i tackle the bottlenecks in planning, review and validation and achieve a constant flow state → tweet
@kunchenguid · 2026-06-07T03:35
claude down again. back to gpt 5.5. i've manually achieved a 99.9%+ high availability setup by making my workflow completely agent-agnostic → tweet
@kunchenguid · 2026-06-06T22:41
asking every employee to burn unlimited tokens is equivalent to encouraging every employee to hire an infinite amount of contractors to work on their behalf. no sane company would ever do that. yet they got successfully brain-washed into thinking tokenmaxxing is a good thing → tweet
@sudoingX · 2026-06-07T17:47
watching an agent code for you feels like: → tweet
AI Models & Performance
@sudoingX · 2026-06-07T17:18
i watched gemma 4 12b build something genuinely impressive today, and then loop itself to death right in front of me. […] so it's not that gemma can't code. it did the hard part beautifully. it cannot finish. it cannot hold a long task together without unravelling […] google has the raw capability, you can see it sitting right there in the code, and then the model loops itself to death on a task a 27b from alibaba finishes clean. → tweet
@sudoingX · 2026-06-07T17:30
here's just the looping part on its own, real time, no edits. gemma 4 12b had already built the thing, the work was done, and instead of stopping it just kept going. the same fragments over and over, talking to itself while the token counter climbed and climbed. watching a model do this is its own special kind of pain → tweet
@sudoingX · 2026-06-07T13:27
looks like gemma 4 12b is trying to kill my rtx 3090. pinned at 100%, 216 watts, 72 degrees […] then you look at the memory bar. 16 gigs. out of 24. and that's at full 256k context. it's not dying, it's barely warmed up. one cheap consumer card, fully local, cooking away with 8 whole gigs to spare. → tweet
@Ex0byt · 2026-06-07T14:45
is there headroom beyond the greedy Babai/GPTQ point for post training quantization…? Or a better transform? Asking for a friend. → tweet
@TheAhmadOsman · 2026-06-07T06:45
Opensource AI MUST WIN. → tweet
@TheAhmadOsman · 2026-06-07T02:26
I got a lot of backlash on this post in January but I knew where we were heading with full conviction so I doubled down like I usually do. I am doubling down again and saying that unless things go catastrophically wrong, Opensource AI is the definite future → tweet
@TheAhmadOsman · 2026-06-07T04:13
Fun LLM Question from Mike today — "Let's say Thanos snaps away every model from existence except Qwen3.5-2B" — What I would do: Phase I: Build synthetic data harness around Qwen3.5-2B. Phase II: Finetune for specialized versions. Phase III: Train a 9B-class model. Phase IV: Raise capital, buy GPUs, scale up. → tweet
Developer Tools & Open Source
@Teknium · 2026-06-07T10:39
You can now change the font in the Hermes Agent dashboard, and independently of the theme.
hermes updateto get access now! → tweet
@Teknium · 2026-06-07T13:24
Way too much stacked for release this week ahh → tweet
@Teknium · 2026-06-07T08:38
We probably need a SOUL Hub don't we haha → tweet
@Teknium · 2026-06-07T08:37
Big PR after reviewing every docs page to ensure comprehensibility, legibility, and accuracy! → tweet
@Teknium · 2026-06-07T08:55
Welcome to the Hermes Agent Contributors Crew! → tweet
@sudoingX · 2026-06-07T08:51
Hermes agent from bangkok traffic. → tweet
@jezell · 2026-06-07T03:40
My Codex /goal to port LibreOffice still going strong after 7 days. Not sure how many weeks this is gonna take, but it's burning a steady billion tokens a day. → tweet
@jezell · 2026-06-07T01:23
Codex really likes to make useless mock based tests. E2E tests with /goal are way better at forcing it to work through problems. → tweet
@jezell · 2026-06-07T01:19
It's really easy to make apps now, but its harder than ever to justify your app existing in the first place. → tweet
@jsuarez · 2026-06-07T15:18
This is Constellation, our lightweight aggregate experiment visualizer. Open source, hackable. Single-file C implementation. Ships with 20,000+ RL experiments worth of data from PufferLib 4. → tweet
@jsuarez · 2026-06-07T15:16
A little perspective: RL as a field spent 10 years making algorithms slower and slower. […] Our whole core realization with PufferLib is that we can write good sims for a lot of problems 10000x faster. Good doesn't even mean accurate. It means accurate enough with domain randomization. […] We ran 20,000 experiments on ~12 GPUs in the 3 weeks leading up to Puffer 4 launch. At traditional speeds, it would have taken Google scale compute. → tweet
@RydMike · 2026-06-07T09:44
Interesting news about #FlutterDev signals package oref, based on alien_signals by same author. Alien signals is probably the fastest signals pkg for Flutter, maybe even the most performant reactive system for Flutter. → tweet
@hnasr · 2026-06-07T14:01
The rule of CPU in the Kernel → tweet
@kunchenguid · 2026-06-07T04:57
is codex the only harness where "exit" is treated as a real prompt and does NOT exit the app? i wonder how much of the world's electricity got wasted because of this → tweet
AI-Assisted Development & DX
@levelsio · 2026-06-07T15:21
🖨️🖼️ Big progress: it can now print images too. I had to install a different printer driver, Claude Code recommended Epson FX-86e which supported images. Then I printed an image, it helped me dump a .bin from the printer raw data output, then it decoded the .bin and figured out how to draw images on a canvas (in the printer) from that data. So crazy I'd have no clue how to do this without AI ever! → tweet
@levelsio · 2026-06-07T12:27
🖨️ It took over a year and lots of help from Claude Code but I have now been able to create a real dot matrix printer on the web that prints from Windows 3.11 in your browser. IT WORKS!! 😊 → tweet
@thdxr · 2026-06-07T17:38
a whole bunch of companies that had good primitives but never figured out DX just got saved by AI. i'm using all these things that were too rough to use before → tweet
@thdxr · 2026-06-07T21:13
liz says when she hears me voice dictating to opencode i'm nicer than i am when talking to people → tweet
@kunchenguid · 2026-06-07T16:01
many people read this as "vibe coded apps are bad so no adoption" […] NO - it's not that simple. adoption of consumer apps depends on distribution. vibe coding doesn't help people solve distribution. with AI, people are going to need less and less apps. why would i switch between 5 different shopping apps if i can do everything in chatgpt? […] app adoption will not just stay flat - it will decrease over time, and increasingly concentrate on a small number of super apps → tweet
Hardware & Infrastructure
@sudoingX · 2026-06-07T13:35
the same grifters who told you to buy a mac mini to run openclaw are quietly selling those machines at a loss now. everyone who bought a used 3090 instead is sitting on a card worth more than they paid for it. compute that does real work appreciates. hype boxes depreciate. → tweet
@TheAhmadOsman · 2026-06-07T20:14
Trying to hire smart engineers? Add GPUs to your job perks and benefits. I know some will sign on the spot if you tell them they get a GB300 Station → tweet
@TheAhmadOsman · 2026-06-07T05:12
Linux monitoring tools you SHOULD have: btop, glances, nvtop/nvitop, duf → tweet
@TheAhmadOsman · 2026-06-07T00:18
Currently working on 4 different articles — LLMs Decoding & Prefilling, LLMs Kernels, CPUs vs GPUs vs Tenstorrent vs Apple Silicon, Differences in "Kernels" → tweet
App Market Dynamics
@levelsio · 2026-06-07T16:29
Ok deployed it on [site] and integrated into the full site. Only thing is perspective is a bit off, so I'll try edit the printer with GPT Image 2 to get a slightly left view of it. But it works, fucking crazyyyyy → tweet