← Tech / AI / IT Monitor Index Tech / AI Generated 2026-06-05 19:31 UTC

Tech / AI / IT Monitor

June 05, 2026 · Based on tweets from the last 24 hours · 141 tweets analyzed · model: ollama-cloud/glm-5.1:cloud

Daily Intelligence Briefing — Tech / AI / IT Monitor

Date: 2026-06-05 | Reporting Window: Last 24 Hours


Executive Summary

NVIDIA's Nemotron 3 Ultra dropped as a major open-source release (550B MoE, 55B active, 1M context), accompanied by full training data and recipes — a significant move for open-weight AI. Google's Gemma 4 QAT weights landed on Ollama, enabling efficient local inference across all model sizes. On the local hardware front, Gemma 4 12B was demonstrated running fully multimodal on a single consumer RTX 3090 at 33 tok/s, underscoring how fast open-source capability is reaching consumer-grade hardware. The agent tooling ecosystem expanded rapidly, with Hermes Agent gaining ElevenLabs integration and desktop improvements, OpenCode shipping parallel git-worktree workflows, and OpenAI rolling out a major ChatGPT memory upgrade with web app publishing. Anthropic's claims about AI-written code ("80% of merged code") faced scrutiny for inconsistencies, while debates about AI company moats and narrative-manufacturing intensified.


Key Events


Analysis

Patterns & Trends:

What to Watch Next: - Whether NVIDIA follows through on continued Nemotron open releases (sm120/sm121 support for RTX PRO 6000/DGX Spark remains a pain point). - How Anthropic responds to the growing scrutiny of its internal AI coding claims and recursive self-improvement narrative. - The competitive dynamics between Duo Agent Platform, Hermes, Cursor, and Codex as pricing pressures mount ($0.25/review fixed price from Duo is aggressive). - Vibe Jam 2026 winners and what they signal about AI-native game development.


Tweet Feed

AI Model Releases & Weights

@TheAhmadOsman · 2026-06-05 00:15

NEMOTRON 3 ULTRA IS HERE MoE - 550B, 55B Active — Base and Instruct in BF16 — NVFP4 Instruct (350GB) — 1M Context Window — Supports MTP Spec Decoding. This release IS NOT just about the Weights It's an Opensource Frontier Intelligence Handbook — Training Data (Pre & Post) — Training Recipes — Environments — Tech Report — Cookbooks Opensource AI FTW → tweet link

@TheAhmadOsman · 2026-06-05 04:48

Finally, today's Nemotron 3 Ultra release makes me very hopeful for the future of Opensource AI. Jensen knows that this is important to keep the powers in check, and I believe he's sincere in his answer to me that there will be continuity to the Nemotron Coalition releases. Big W → tweet link

@ollama · 2026-06-05 18:32

Gemma 4 Quantization-Aware Training (QAT) weights are now available on Ollama! They reduce memory requirements while maintaining model quality. E2B: ollama run gemma4:e2b-it-qat / E4B: ollama run gemma4:e4b-it-qat / 12B: ollama run gemma4:12b-it-qat / 26B: ollama run gemma4:26b-a4b-it-qat / 31B: ollama run gemma4:31b-it-qat → tweet link

@ollama · 2026-06-04 23:34

ollama run gemma4:12b — Gemma 4 12B is updated on Ollama, and available across all platforms! Try it on: Claude Code / Hermes Agent / OpenClaw / Codex / Codex App → tweet link

@victormustar · 2026-06-05 14:18

Nemotron 3 Ultra is on HuggingChat ☘️ its speed/performance ratio makes it the best model right now for daily chats imo + you get @ExaAILabs available at no extra cost! → tweet link

@SkylerMiao7 · 2026-06-05 02:53

MiniMax-M3, best open-source model right now → tweet link

@victormustar · 2026-06-05 08:21

RT @osanseviero: Introducing Magenta RealTime 2 🎺 — Open model for live music generation, Just 2.4B parameters, perfect for on-device → tweet link

@victormustar · 2026-06-04 22:53

RT @boson_ai: Higgs Audio v3 TTS is here. Built for voice AI that speaks, not just reads: 100 languages with single-digit WER/CER → tweet link

@victormustar · 2026-06-04 19:00

RT @PiotrZelasko: Second big release from us today: Nemotron-3.5-ASR-Streaming! 🌎40 languages ⚡️80ms - 1s controllable latency 🔥240 - 2400 → tweet link

Local AI & Consumer Hardware

@sudoingX · 2026-06-05 12:52

a year ago this wasn't possible on a single consumer gpu. today it just is, and i'm still a little stunned by it. google dropped gemma 4 12b two days ago and i loaded it on one 3090. natively multimodal, text image video audio all in one net, no separate encoders bolted on, 256k context, apache 2.0, weights i actually own and can run offline forever. [...] open source is so back, and it's moving faster than the labs are pricing in. → tweet link

@sudoingX · 2026-06-05 13:20

watch gemma 4 12b q8 dancing on a single rtx 3090 at 33 tokens a second average. google dropped this two days ago and it's the kind of thing that quietly moves the floor. a fully multimodal model, text image and audio in one net, 256k context, apache licensed, running entirely on one consumer gpu, no one metering your tokens. [...] a year ago this needed someone else's datacenter. today it's a card you can buy. open source isn't catching up anymore, it's setting the pace. → tweet link

@TheAhmadOsman · 2026-06-05 04:34

All it takes to get started with Local AI is a single RTX 3090, so go buy that GPU anon → tweet link

@TheAhmadOsman · 2026-06-05 03:05

Guys, just a reminder that WE ALL have to speak up and ask NVIDIA to put more effort & focus on the sm120 and sm121. I am just like you annoyed with this situation for the RTX PRO 6000s and DGX Sparks. But I also believe if we all voice this they'll hear us & do the right thing → tweet link

@badlogicgames · 2026-06-04 23:23

i lied. pibot is now multi-user, serving each kid in the hood from a single m1 max with parakeet for stt, gemma 4 24b a3b as the llm via llama.cpp, and qwen3-tts for tts. it's awesome! maybe tomorrow we'll have our robot building session. → tweet link

Agent Frameworks & Developer Tools

@Teknium · 2026-06-04 23:15

Welcome to the Hermes Crew, ElevenLabs! → tweet link

@Teknium · 2026-06-05 17:43

The desktop app now has a Chinese language support! → tweet link

@Teknium · 2026-06-05 17:34

Another banger by Tonbi - Hermes Agent masterclass series video on model options and configurations! → tweet link

@Teknium · 2026-06-05 17:34

We had to make some deep level changes to Hermes Update command this morning. It may require a number of you to run hermes update twice in a row (where you'll see an error the first time). Please run it twice and you should be good to go. → tweet link

@Teknium · 2026-06-05 09:03

I exclusively build Hermes Agent with Hermes Agent as well! → tweet link

@Teknium · 2026-06-05 09:33

Would love to make the plugin developer experience better! Please let me know how I can help. → tweet link

@Teknium · 2026-06-05 01:50

Some more clarity and updates on this doc to hopefully get you sorted. The desktop app is in preview and we are working hard to make these kinds of use cases more seemless and robust. Basic auth is now required and oauth options are coming! → tweet link

@thdxr · 2026-06-05 00:39

we landed on a pretty good workflow for doing parallel work in OpenCode. this demo is with git worktrees but i also preview an alternative we're working on at the end. this will be in 1.6.0 → tweet link

@thdxr · 2026-06-05 04:47

kit has been building a discord bot with opencode 2.0 to test it. i'm actually loving it mostly for the collaboration. it's very fun and productive to prompt the bot as a team. it even makes talking to each other easier because we can pull in opencode to look stuff up and explain things → tweet link

@thdxr · 2026-06-05 13:18

even if i'm not actively shipping code i'm constantly talking to opencode to think through ideas. you can see on days where i'm busy with company stuff my token count goes down. so unfortunately this is a good measure of something → tweet link

@TheAhmadOsman · 2026-06-05 01:04

Let me make your Codex Cli experience better with — Multi-agent delegation — Enhanced memory — Better artifacts — Children AGENTS.md contextualization — Runtime metrics (optional) → tweet link

@kunchenguid · 2026-06-05 15:57

my codex profile. longest task is not accurate because i do gnhf which orchestrates fresh context windows in a loop for long running tasks. everything else checks out and i've been really loving gpt 5.5 fast mode which is my default now → tweet link

@MengTo · 2026-06-05 10:09

People ask what I use for screen recording. I built one in Codex because I was tired of jumping between Screen Studio and CapCut every time I made content. [...] That's what I love about AI: you can build the tool you wish existed. → tweet link

@steipete · 2026-06-05 13:28

RT @ChrissGPT: Wait, this is actually a goated memory update, why isn't anyone talking about this? Claude has "Dreams" in its agent/API st… → tweet link

@steipete · 2026-06-05 13:22

RT @OpenAIDevs: More of the iOS app loop, now inside Codex. The Build iOS Apps plugin lets Codex view and test your iOS app in the in-app… → tweet link

Inference Infrastructure & Tooling Deep Dives

@TheAhmadOsman · 2026-06-04 21:21

Everything You Need To Know About Inference Engines and Running LLMs Locally at Home — Explains why Inference Engines exist in the first place: Prefill is not Decode / VRAM is not bandwidth / Fit is not speed / KV Cache is the real memory problem / Quantization only matters if the engine has good kernels / MoE and the routing problem / How long context changes the serving problem / Multi-GPU changes the interconnect problem [...] Maps the Engines: llama.cpp, MLX, ExLlamaV3, vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo → tweet link

@alexinexxx · 2026-06-04 19:07

RT @vikhyatk: Wrote a post about how Photon (Moondream's inference engine) hides GPU bubbles using pipelined decoding. Speeding up inference… → tweet link

@jsuarez · 2026-06-05 00:11

That's Puffer Constellation, a lightweight experiment visualizer in C & raylib. View our ~20,000 baselines for 4.0 live in your browser → tweet link

OpenAI & ChatGPT

@sama · 2026-06-04 22:17

big upgrade to chatgpt memory rolling out today! → tweet link

@sama · 2026-06-04 22:21

build and publish web apps with chatgpt! i really wish i had this when i was a kid, but i do miss hypercard. → tweet link

Anthropic Scrutiny & Industry Commentary

@kunchenguid · 2026-06-05 03:31

hmm there's an inconsistency in what anthropic is telling the world. tweet below says "80% of code is written by claude". but in a recent talk, they told us "there's no more manually written code anywhere at the company". so which is it? is the 80% number even trustworthy? this seems to validate what i shared in my recent tokenmaxxing explainer — it's a narrative manufactured to make other companies feel behind and want to burn more tokens → tweet link

@kunchenguid · 2026-06-04 20:29

as an industry, we really need to move away from talking about volume of code. anthropic writing 8x more code means absolutely nothing. only thing matters is the outcome. how about let's look at fixing claude uptime? → tweet link

@kunchenguid · 2026-06-05 02:24

this is hilarious! everybody SLOW DOWN and let me be the only one that keeps charging ahead with my "confidential" IPO which i tweeted about → tweet link

@TheAhmadOsman · 2026-06-05 13:09

I am not saying that Anthropic is a cult. But if they were a cult, what would they do any differently? lol → tweet link

@TheAhmadOsman · 2026-06-05 11:10

Anthropic is ngmi unlike what you all might think btw → tweet link

@TheAhmadOsman · 2026-06-04 21:51

Dario when he realizes Opensource AI is catching up and Anthropic has no moat + last Christmas Claude Code boom is never going to happen again → tweet link

@TheAhmadOsman · 2026-06-05 03:38

Funniest thing would be Anthropic getting nationalized on the 1 year anniversary of this tweet 🤡 → tweet link

@TheAhmadOsman · 2026-06-04 20:05

When a system mediates what you read, what you write, what you remember, and what options you even notice, understanding that system stops being a hobby and becomes literacy → tweet link

@TrungTPhan · 2026-06-05 16:16

Sarah Connor seeing the "recursive self-improvement" animation on Anthropic's "When AI builds itself" blog post → tweet link

@thdxr · 2026-06-05 18:59

we've seen reports of this for a while now — almost a year. they investigated and concluded they're very convincing hallucinations. super weird because you get very specific, very random responses → tweet link

Competitive Landscape & Evaluations

@RealGeneKim · 2026-06-05 14:06

RT @bstaples: I expected Duo Agent Platform to beat Cursor, Claude, Copilot and Devin on price with our $0.25/review fixed price, I did not… → tweet link

@louszbd · 2026-06-05 07:40

RT @arena: Introducing Agent Arena: real-world agentic evals at scale. How do you evaluate agents doing actual work? We measure millions o… → tweet link

@jezell · 2026-06-04 19:48

RT @eastdakota: Welp, that happened faster than I predicted. Thought it would be end of 2027, then early 2027, but agentic traffic growing… → tweet link

@jezell · 2026-06-04 21:38

I remember when Google did this and a couple days later Apple made them take it out. Wonder if OpenAI will fair better. → tweet link

AI Music & Creative Applications

@LinusEkenstam · 2026-06-05 08:09

I still can't believe we walked into a living litmus test of AI music in our Uber back from the bar in SF. The driver had no clue that his current favorite music that he was blasting was 100% AI generated. completely oblivious to the fact. [...] After some contemplation, he said, I still love it and it's great. suno is def worth more than $5bn → tweet link

Open Source Ecosystem & SaaS Disruption

@FinansowyUmysl · 2026-06-05 07:25

To jest to czego zazwyczaj piszący tu programiści nie rozumieją. Kosztem niedopasowanego SaaS-a nie jest tylko jego koszt subskrypcji, ale również koszt alternatywny w postaci straty czasu i nieefektywności procesów [...] A dzięki viebecoding koszt tworzenia znacznie spadł → tweet link

@iamdevloper · 2026-06-05 11:09

startups will create a "war room" and it's just 3 guys looking at some AWS logs → tweet link

@levelsio · 2026-06-04 21:14

🏆 Round 2 of judging the Vibe Jam of 2026 sponsored by @cursor_ai + @boltdotnew + @heyglif + @tripoai is finished now. After going through 945 submissions in Round 1, me and @s13k_ then rated the top 100+ in Round 2 [...] Top 25 go to Round 3. Prizes: 🥇 $25,000 🥈 $10,000 🥉 $5,000 → tweet link

@levelsio · 2026-06-05 13:39

Also in little cool achievements: 🕹️ Vibe Jam 2026 games have now been played by over 1,000,000 people. And we've got almost 50,000,000 impressions on X! → tweet link

@steipete · 2026-06-05 15:05

RT @charliermarsh: Ladybird is no longer accepting public pull requests. I don't know what to do about it yet, but the dynamics of open so… → tweet link

@steipete · 2026-06-05 12:35

RT @hungv47: It's a crime that this repo only has under 400 stars. /autoreview and /handoff are literally two of the best agent skills I'v… → tweet link

@Teknium · 2026-06-05 01:46

RT @cyrilXBT: 140K GitHub stars. Three months. Most people still copy-pasting into ChatGPT every morning. Hermes plus Obsidian plus Notebo… → tweet link

Developer Insights & Engineering Culture

@thdxr · 2026-06-05 01:36

how to be good at your job — realize this one thing is actually made up of two separate things — realize instead of solving the direct problem you can solve a broader problem — instead of implementing thing, implement other thing that makes it easier to implement thing → tweet link

@jsuarez · 2026-06-04 23:38

This post right here officer. Let me know when your engineers ship 8x LESS code → tweet link

@jsuarez · 2026-06-05 18:58

Reinforcement learning research with Joseph Suarez → tweet link

@TheAhmadOsman · 2026-06-05 07:02

RT @TheAhmadOsman: Step-By-Step LLM Engineering Projects Roadmap — Build a tokenizer — Learn embeddings — Implement RoPE / ALiBi — Hand-wi… → tweet link

@TheAhmadOsman · 2026-06-05 06:50

RT @TheAhmadOsman: Local AI hardware = capacity × bandwidth × software stack — Capacity tells you what fits — Bandwidth tells you how hard… → tweet link