← Tech / AI / IT Monitor Index Tech / AI Generated 2026-06-06 19:31 UTC

Tech / AI / IT Monitor

June 06, 2026 · Based on tweets from the last 24 hours · 127 tweets analyzed · model: ollama-cloud/glm-5.1:cloud

Executive Summary

The past 24 hours saw an extraordinary burst of open-weight AI model releases—over 25 notable drops across LLMs, image gen, audio, and video—headlined by NVIDIA's 550B Nemotron 3 Ultra, Google's Gemma 4 12B with QAT, and Ideogram 4's first-ever open weights. Meanwhile, Nous Research shipped Hermes Agent v0.16.0 with a full Desktop GUI app and dashboard overhaul, rapidly emerging as a preferred agentic orchestration layer. On the hardware front, NVIDIA's DGX Spark proved itself as a legitimate from-scratch research machine, while the community pressed for real benchmarks on AMD's Strix Halo competitor. A widely-cited poll confirmed that 75% of developers now review code written by agents rather than writing it themselves—the agent era is no longer emerging, it is the median workflow for those closest to it.

Key Events

Analysis

Patterns: This week marks a clear inflection point for open-weight AI. The sheer volume—25+ releases across LLMs, vision, audio, video, and 3D—suggests the open ecosystem has reached critical mass and is no longer lagging frontier closed models by years, but by months. The quality of smaller models (Gemma 4 12B, Liquid AI LFM2.5-8B, JetBrains Mellum2-12B) means consumer-grade hardware is now sufficient for serious work.

Escalation: The developer workflow is shifting faster than most commentary acknowledges. The 75% agent-coding poll and reports of developers feeling "guilty" leaving agents unsupervised suggest the human role has already moved from author to reviewer/operator. Agentic orchestration layers (Hermes, Codex, Claude Code) are competing fiercely, with reliability and tool-calling quality becoming the key differentiator—not raw model capability.

De-escalation: Frustration with current agentic tools (Claude Code unreliability, Codex rate limits, Grok Build feature gaps when switching models) shows the "review seat" workflow still has sharp edges. Local AI hardware remains fragmented— DGX Spark has real results while AMD Strix Halo has hype but no posted benchmarks.

What to watch next: Whether AMD Strix Halo produces real community benchmarks; whether Hermes Agent's local model support at small scales (3B-12B) sustains multi-step agentic reliability; NVIDIA's response to sm120/sm121 issues; and whether the open-model avalanche triggers a defensive move from closed-weight providers on pricing or capability.

Tweet Feed

Open-Weight AI Model Releases

@victormustar · 2026-06-05T21:59

Before the week ends, let's acknowledge one of the most INSANE week ever for open AI, with 25+ notable open-weight drops across every modality: 🧠 LLMs → NVIDIA Nemotron 3 Ultra: 550B hybrid Mamba-MoE, only 55B active, 1M context, MMLU 89.1. NVFP4 variant claims ~5x throughput on Blackwell. First openly-weighted 550B hybrid Mamba-Transformer, closing the gap with frontier closed models. → Google Gemma 4 12B: fully open dense any-to-any (text/image/audio/video), 256k context, encoder-free, 140+ languages, AIME 2026 at 77.5. Shipped with a 23-checkpoint QAT wave (mobile ONNX + MLX). Most deployable model of the week. → StepFun Step-3.7-Flash: 198B sparse MoE VLM, ~11B active, SWE-Bench PRO 56.3. Apache 2.0. → Liquid AI LFM2.5-8B-A1B: edge MoE, just 1.5B active, 128k ctx, MATH500 88.8, MLX-ready. Best on-device option this week. → JetBrains Mellum2-12B-A2.5B-Thinking: their first open MoE, near-Qwen3-14B coding at 2.5B active. Apache 2.0. 🎨 Image gen (the surprise of the week) → Ideogram 4: their FIRST-EVER open weights. 9.3B flow-matching DiT trained from scratch. #2 overall behind GPT Image 2, top open-weight model on Design Arena + LMArena. Strongest open checkpoint for text-rich images, full stop. It has taste. Still can't believe this is open weights. 🔊 Audio & Speech (a breakout week for open TTS, 4 labs shipped) → Boson Higgs Audio v3 4B: 102 languages, 21 emotions, singing/whispering/shouting, sub-second TTFA. → RedNote dots.tts: the only fully continuous (no codec) open TTS pipeline, Apache 2.0. → Google Magenta RealTime 2: real-time music gen, <200ms latency, text+audio+MIDI. multimodalart ported it to PyTorch within hours with live ZeroGPU demos. → NVIDIA Nemotron-3.5 ASR: 600M streaming, 17x more concurrent streams vs Parakeet RNNT 1.1B. 👁️ Vision & VLMs → PaddleOCR-VL-1.6: SOTA document parsing at 1B params, Apache 2.0. → Baidu NAVA: 6.3B joint audio-video gen, best-in-class A/V sync, Apache 2.0. 🎬 Video, 3D & World Models → NVIDIA Cosmos3-Super: 64B omnimodal world model coupling action trajectories with video+audio gen, for Physical AI. → JD JoyAI-Echo: up to 5-min multi-shot text-to-video on LTX-2.3. → ByteDance Bernini-R + VAST TripoSplat (single-image-to-3D Gaussian splats, MIT). → tweet link

@TheAhmadOsman · 2026-06-05T20:15

Great news. Google just released the QAT (4bit) of their Gemma 4 model series including the 31B Dense and the 26B MoE. Another W for Opensource AI this week → tweet link

@TrungTPhan · 2026-06-06T18:08 (RT @bearlyai)

[NEW] The Bearly AI app just updated its video-generating feature. Try Kling 3.0 Omni (along with access to Gemini, ChatGPT… → tweet link


Hermes Agent v0.16.0 & Nous Research Ecosystem

@Teknium · 2026-06-06T01:49

Hermes v 0.16.0 is out now! This release includes all the updates you've heard about this week and more! - The Desktop GUI App - The Overhaul to the Dashboard - Leaner Built-In Skillset - New Security Layers for Remote Dashboard & GUI Access - and much more! → tweet link

@Teknium · 2026-06-06T04:47 (RT @iamlukethedev)

Hermes Agent v0.16.0 (2026.6.5) just dropped. This is arguably the biggest Hermes release yet. Hermes is no longer just… → tweet link

@Teknium · 2026-06-06T09:49

Updated the skills hub integration search in the Hermes Agent Dashboard to be way more comprehensive and give you the info you need. → tweet link

@Teknium · 2026-06-06T04:46 (RT @NousResearch)

Hermes Desktop 现已支持简体中文——聊天界面完整适配简体中文。桌面应用现在已在所有 UI 界面全面提供简体中文翻译 → tweet link

@Teknium · 2026-06-05T21:12 (RT @akshay_pachaar)

i just built a 4-agent software team. everything runs from Telegram and gets managed on a kanban board. a project man… → tweet link

@Teknium · 2026-06-06T05:46 (RT @akshay_pachaar)

The Hermes Desktop App is insanely good. It's now the best way to run AI agents on your computer. Here's the full set… → tweet link

@Teknium · 2026-06-06T00:21 (RT @imbabybrooklyn)

Profiles in @NousResearch Hermes Desktop. Run sessions across profiles at once, switch between them mid-thread, and dra… → tweet link

@Teknium · 2026-06-06T13:40 (RT @_avichawla)

The top Hermes integrations to give your agent superpowers: 1. Obsidian — It works as a Karpathy-style second brain, but on… → tweet link

@louszbd · 2026-06-06T06:57 (RT @NousResearch)

Hermes Desktop 现已支持简体中文 → tweet link

@gospaceport · 2026-06-06T03:02 (RT @Teknium)

We had to make some deep level changes to Hermes Update command this morning. It may require a number of you to run hermes upd… → tweet link


Local AI Hardware: DGX Spark, RTX PRO 6000, AMD Strix Halo

@sudoingX · 2026-06-06T18:56

people keep asking what i actually do with the dgx spark sitting on my desk. i kept answering with inference benchmarks, large models running reliably, big context, the usual. and that's all real, but it was never the point. it's not an inference box. it's the smallest research lab i've ever owned. so here's the one thing. i'm training a neural net from scratch to find exoplanets in starlight. no pretrained weights, no borrowed architecture, my own design, on real Kepler data. right now it sits at 0.75 on a tiny slice, 514 stars out of 7,586. proof of life, nothing more. now the move is to go big. first the full labeled archive, fifteen times the data. then the part i'm itching toward, self-supervised pretraining on around 200,000 raw light curves, a corner of machine learning i've genuinely never worked in before. and this is where the box stops being a gimmick. one gb10, 121gb of unified memory, so the entire dataset lives in ram and i train for days with zero data-loader bottleneck, the thing that usually kills experiments like this on a normal rig. no streaming, no thrashing, just the whole sky sitting in memory. everyone treats the spark as an inference box. on my desk it's a from-scratch research lab. → tweet link

@sudoingX · 2026-06-06T13:02

the more i use my dgx spark the more i think it's one of the most undervalued machines on the market right now. and i keep finding new things to throw at it that have no business working on something this small. but every time i post about it the same question comes back, what about the amd strix halo. and here's what actually bugs me, i can't answer it, because i've never run one myself, or even seen anyone around me run one. tons of people name it as the competitor, almost nobody posts real numbers. no tok/s, no model loads, no thermals, nothing. just the name. so i'm asking straight up. if you've got a strix halo or a ryzen ai max box, drop your real numbers. what models, what speeds, what breaks. is it actually competing with the spark, or is it the machine everyone recommends and nobody runs. → tweet link

@TheAhmadOsman · 2026-06-06T13:52

Folks at NVIDIA are hiring for someone to help them fix sm120 and sm121 for Local AI btw. Good opportunity if you're passionate about Local AI → tweet link

@TheAhmadOsman · 2026-06-05T19:57

They hear your concerns on the sm120 and sm121. Whoever is running this NVIDIA account deserves a raise btw 💚 → tweet link

@TheAhmadOsman · 2026-06-06T01:04

Gerardo is a Senior Director of Product at NVIDIA AI. Reply to him with your RTX PRO 6000 / DGX Spark issues and I am confident he'll forward them to the right people to take care of → tweet link

@TheAhmadOsman · 2026-06-06T01:34

The other day someone asked me to if they should sell their RTX 3090 after purchasing an RTX PRO 6000. I told them I wouldn't, and I said that for a reason. Disclaimer: Not a financial advice → tweet link

@TheAhmadOsman · 2026-06-06T07:18

This is why you cannot afford a 32GB RAM kit anymore and it's freaking beautiful → tweet link

@gospaceport · 2026-06-05T22:44

96GB gets an agent VERY FAR now! I wouldn't chase the edge of the frontier locally (most shouldnt either imo) as that now crest the 1T range, not feasible (and frankly that needed outside developer tasks now. The entire down model range however has gotten crazy good for us thankfully! FWIW I am undoing my 8x GPU rig, I only use 4x GPU's at a time most of the time and I miss having it on set. I am also now married to Qwen 3.6 27b at full precision. Killer model for non developer-centric tasks. → tweet link


AI-Native Developer Workflows & Agent Coding

@sudoingX · 2026-06-06T09:05

i posted this half asleep at 4am expecting maybe a hundred votes and a laugh. 278 people answered, and three out of four said they just review their code now. sit with that for a second. not "agents help sometimes," not "i use copilot for autocomplete." three quarters of the builders i reached have handed the actual writing to the machine and moved into the review seat. only 1 in 20 still writes it all by hand. there was no announcement, no headline, no single day it flipped. it just quietly became true while everyone was still arguing about whether it would. this isn't the start of the agent era, it's already the median workflow for the people living closest to it, and most of the world hasn't noticed because the ones living it are too busy shipping to write thinkpieces about it. we're not at the beginning of this. we're in the middle of it. → tweet link

@sudoingX · 2026-06-06T12:14

i was genuinely thrilled when composer 2.5 landed in grok build, right up until switching to it killed the one feature i actually use grok build for. the reason i live in grok build is the native X access. ask it about a post, an account, a whole thread, and it pulls it live, no bot walls, no scraping. on a grok model that x_search tool is right there and it's incredible. switch to composer 2.5 in the exact same harness and it's just gone. it falls back to a generic web fetch, hits X's login wall, sees nothing. composer even explains it itself, the x_search tool isn't wired to it. that's the part that gets me. same cli, same harness, but the feature i came for disappears the moment i pick the better coding model. as a user i don't expect that. if grok build ships composer 2.5, the X-native tools should come with it. this reads like a wiring gap, not a hard limit. close it and composer 2.5 plus grok build's live X access is the best agentic coding setup on the platform, full stop. right now i'm forced to choose. xai bros please don't make me choose. → tweet link

@sudoingX · 2026-06-06T09:41

dario. wtf dude. this run was nonnegotiable. claude code is so fucking unreliable man. → tweet link

@alexinexxx · 2026-06-06T00:03

rate-limited by codex, lied to by claude, barely holding it together → tweet link

@gdb · 2026-06-06T03:35

so much more fun to use a computer via codex → tweet link

@nummanali · 2026-06-06T00:24

TIL the Codex remote feature works even if it's not your account if you use SSH login. Super useful if you're using multiple accounts or like me, assisting friends and family. Setup Tailscale, login with the device password and bobs your uncle → tweet link

@sudoingX · 2026-06-05T21:22

it's 4:30am and i just caught myself feeling guilty for going to sleep and leaving my agents running without me. sit with that for a second. i felt bad leaving software alone to work, like they're coworkers pulling the night shift and i'm the one clocking out early. what a time to be building. → tweet link

@sudoingX · 2026-06-05T20:21

to be clear on why i come down hard on openclaw, it's not really the tool, it's that for a lot of people "agent" has quietly come to mean openclaw, like it's the only option. it's not even close. hermes agent is my main orchestration. it's a beast at coding, i trust it with real work, and paired with opus 4.8 on max effort it's the most capable agentic workflow i've ever run. → tweet link

@TheAhmadOsman · 2026-06-06T13:03

Best thing I ever done was setup a telegram bot on an RTX 3070 to handle all my wife's media requests via Plex. I am no longer in charge of fixing everything that bothers her about it and she doesn't have to wait for me. Bonus: makes excellent recommendations according to her → tweet link

@hnasr · 2026-06-06T14:10

The thing AI can't give software engineers → tweet link

@TheAhmadOsman · 2026-06-05T21:20

Intuition and taste are cultivated through time by grinding through the fundamentals and understanding how every piece fits in the puzzle. Taking shortcuts will not get you there, and if you make a win it will most probably be shortlived. Life ain't vibecoding → tweet link

@gdb · 2026-06-06T00:31

email integration with chatgpt → tweet link

@jezell · 2026-06-05T23:53 (RT @colemurray)

orgs deploying background agents aren't focusing on agent optimization (yet). across deployments with OpenInspect, companie… → tweet link

@steipete · 2026-06-06T13:42 (RT @jenzhusott)

Massive output uptick due to agentic AI. Complete flat adoption. → tweet link


AI Inference Engineering: Kernels & LLM Decoding

@TheAhmadOsman · 2026-06-06T02:17

You don't "run a model" — You run Kernels. The model is just a graph. The Inference Engine is scheduler / optimizer / executor. But the actual work? That happens in the Kernels — MatMul Kernels, Attention Kernels, RMSNorm Kernels, KV cache Kernels, Quantized linear Kernels, Sampling Kernels, Fused "please don't write this back to memory 9 times" Kernels. Same model, same GPU, same VRAM — Wildly different performance. Because one stack is using optimized fused Kernels that understand your hardware, and the other stack is playing hot potato with tensors through 47 tiny launches and pretending the GPU is the problem. Bad Kernels make people say: "this model is slow". Good Kernels make people say: "wait how is this running locally?" This is why Inference Engines and the Kernels implemented within them matter. The model is the recipe. The hardware is the kitchen. The Kernels are the knives, pans, burners, and the chef not cutting onions with a spoon. Most people benchmark models. The real ones benchmark the Kernels underneath. → tweet link

@TheAhmadOsman · 2026-06-06T09:49

LLM Decoding Simplified. From the upcoming article on X → tweet link

@TheAhmadOsman · 2026-06-06T07:41

LLM inference is mostly about avoiding unnecessary movements btw → tweet link


Local Model Deployment & Agentic Testing

@sudoingX · 2026-06-05T20:47

benchmarks were the easy part. speed, vram, context, i already have those. the question that actually matters is the one nobody screenshots: is this thing actually agentic. so here's the setup. google's brand new gemma 4 12b dense, running fully local on a single 3090, q8 so the precision is all there, pointed straight at hermes agent. now we find out if a 12b can hold a real multi-step task, call tools without fumbling, write files, recover when it's wrong, or whether it falls apart the second you ask it to actually do something. → tweet link

@sudoingX · 2026-06-05T21:17

watch how easy it actually is to point hermes agent at a local model. you drop in a local http url, it auto-detects the model, and you're talking to it. no provider configs to fight, no yaml hell, no digging through docs for the right format. the url goes in and gemma 4 12b is just running on my own 3090, smooth like butter the whole way. → tweet link

@sudoingX · 2026-06-05T19:59

small local model that falls apart in bloated agents like openclaw just runs like a wild horse in hermes agent. and that's not even my line, someone else called it that, i've just been quietly pointing people at this harness for months because it held up on everything i threw at it, 3b models all the way to one trillion params. watch this happen on my own machine. i pointed hermes agent at a local http endpoint, gemma 4 12b on my 3090 llama.cpp server, and it auto-detected the model and started working immediately. no config wrestling, no broken tool calls, no babysitting the output format, i typed in a url and it just went. → tweet link

@sudoingX · 2026-06-06T07:46 (RT @sudoingX)

small local model that falls apart in bloated agents like openclaw just runs like a wild horse in hermes agent. → tweet link

@sudoingX · 2026-06-05T21:23 (RT @sudoingX)

watch gemma 4 12b q8 dancing on a single rtx 3090 at 33 tokens a second average. google dropped this two days ago and it's t… → tweet link


ML Research: Exoplanet Detection from Scratch

@sudoingX · 2026-06-06T13:34

a few days ago i posted this dip and said a net i built from scratch found a planet in it. fair question came back, ok, but can it actually tell you anything real about that planet. so i taught it the next step. take that same dip and pull the planet's actual physics out of it. not just detect, measure. here's what came out, next to what NASA has on file. radius, orbit distance, temperature, all derived from four years of public starlight on my own hardware. every number within a percent or two of the archive, the orbit distance almost exact 0.0472 AU against their 0.0473. → tweet link

@sudoingX · 2026-06-06T15:47

and here's the part that still messes with me. this is the orbit i pulled out of that dip. the solid loop is Kepler-8b going around its own star. the dashed one is our Mercury, dropped in as a yardstick you already know. and Kepler-8b's whole orbit fits about eight times inside Mercury's. and that's the orbit that's tiny, not the planet, Kepler-8b itself is a gas giant bigger than Jupiter, just whipping around on a brutally short leash. it's not falling in, it's a stable orbit, just an extreme one. a complete alien world reconstructed from a 0.9% flicker in starlight recorded over a decade ago. → tweet link


Developer Tools & Frameworks

@jezell · 2026-06-05T23:24 (RT @binsquares)

smolvm has hit stable release: v1.0.0! You can now fork smolvm. It means you can fork to create virtual machines off of a… → tweet link

@jezell · 2026-06-05T20:29 (RT @vintcessun)

AI agent 要跑沙箱隔离,但冷启动微VM太慢了。这个项目干脆把预热好的父 VM snapshot 当进程 fork:子 VM 共享内存直到写时才复制,100 个 KVM 隔离的沙箱 100ms 就能全出来。相当于把 fork 的开销套在 VM上 → tweet link

@thdxr · 2026-06-05T23:32 (RT @jlongster)

a quick recording of how I approach dealing with layout shifts → tweet link

@RydMike · 2026-06-05T23:17

A fix for this "dart format ." issue when using #FlutterDev packages that use Swift Package Manager on iOS and macOS, is being worked on, thank you @dart_lang and @FlutterDev 🙏🙂💙 → tweet link

@RydMike · 2026-06-06T03:15 (RT @algebrandon)

Coming to Flutter Scene 0.16.0. 🫖 → tweet link

@RydMike · 2026-06-06T03:34 (RT @ulusoyapps)

📢 I shared a new article on GenUI 🚀 This post sets up the next one, where I will look at Flutter's… → tweet link

@RydMike · 2026-06-06T03:17 (RT @divyanshub024)

Flutter desktop apps are underrated 🔥 Been building something cool with Flutter for desktop lately → tweet link

@RydMike · 2026-06-05T23:34 (RT @MishaalRahman)

🐤 2606 Android Canary is now available with an exciting new experimental feature: more theming options! → tweet link

@kunchenguid · 2026-06-06T15:07 (RT @petergyang)

My next guest, @kunchenguid was an L8 principal engineer at Meta and Microsoft who recently quit to build AI products solo… → tweet link


AI Industry, Startups & Impact

@sama · 2026-06-06T01:56 (RT @SavinovNikolay)

Excited to share that I've joined OpenAI in London to work on pretraining! I've spent the last few years on pretrainin… → tweet link

@TrungTPhan · 2026-06-06T15:17 (RT @bearlyai)

Stack Overflow has seen the number of monthly questions on its platform collapse from 300k to ~0 since launch of ChatGPT. → tweet link

@levelsio · 2026-06-06T09:42

Yes but SEO suffers from the same issue. Before AI, you'd be competing a few competitors who put time and effort into writing pages that ranked on Google. After AI, you're now competing with 1000x more because anyone can generate pretty human sounding pages that rank on Google → tweet link

@levelsio · 2026-06-06T08:59

I think the challenge is that everyone can now build apps. But 1) almost nobody has distribution (like an audience), or 2) the money to pay for distribution (ads or UGC), or 3) the creative genius to get distribution for free (classically called guerilla marketing) → tweet link

@levelsio · 2026-06-06T09:04

Logically how this ends is you will have 3 groups being successful in startups: 1) VC funded startups with money can still get distribution by ads/UGC and all their money will go there (because building is cheap with AI) 2) influencers will become more important 3) you still get a few creative geniuses who can hack going viral → tweet link

@levelsio · 2026-06-06T10:28

Great post by @smalzner, the founder of Franz. He launched 10 years ago, went superviral, got lots of offers to get VC funded, didn't see the point. So he fired everyone and went back to a team of one. Just him (and probably AI) shipping very very fast! → tweet link

@TheAhmadOsman · 2026-06-06T03:39

I wanna teach a course on LLMs 101 in an educational institution + have it recorded and available online to the public for free. I still have an email thanking David J. Malan when I was 13 for putting CS 50 online for free for me to watch it in Egypt. Who knows maybe it'll happen → tweet link

@steipete · 2026-06-05T20:43 (RT @ycombinator)

We're excited to announce Peter Steinberger as a speaker at Startup School 2026! @steipete is the creator of OpenClaw, th… → tweet link

@steipete · 2026-06-06T06:25 (RT @cnakazawa)

The answer to receiving more contributions with agents is not going closed source, it's using more agents to maintain it. → tweet link

@steipete · 2026-06-06T06:23 (RT @shanselman)

VibeOS - Fully Hallucinated Operating System from Microsoft BUILD #msbuild by @stevensanderson (genius) (relax, it's a joke) → tweet link

@thdxr · 2026-06-06T14:24

i've been thinking about this since i first tried a waymo. someone can confidently stand in front of it and fuck with you however they want and you're just stuck. can't do this with a real driver because you don't know they won't run you over → tweet link

@jsuarez · 2026-06-05T21:53

In PufferLib, we just ignore PRs that are not also brought up in the discord or dev streams → tweet link

@jsuarez · 2026-06-05T20:12 (RT @rosinality)

Training an RNN in parallel by RNN cell to predict a compressed state that predicts future outputs… → tweet link