← Tech / AI / IT Monitor Index Tech / AI Generated 2026-06-29 19:31 UTC

Tech / AI / IT Monitor

June 29, 2026 · Based on tweets from the last 24 hours · 140 tweets analyzed · model: ollama-cloud/glm-5.1:cloud

Executive Summary

NVIDIA's largest-ever model, Nemotron 3 Ultra (550B parameters, hybrid Mamba Transformer MoE), failed a real-world agentic coding test—running 100 minutes and burning 440K tokens only to declare success on a blank screen, exposing critical gaps in long-horizon agent verification. Meanwhile, the Local AI movement accelerated with a dedicated summit at AI Engineer World Fair, NVIDIA opening free hosted access to its model lineup, and Framework Desktop demonstrating local inference gains. The developer tooling ecosystem expanded rapidly: Cursor launched iOS with cloud agents, OpenCode v2 added subagent management, and Hermes Agent rolled out voice integration and Stripe capabilities. A notable shift emerged toward "Return on Tokens" (ROT) as the industry narrative pivots from token maximization to token optimization.

Key Events

Analysis

Pattern: Agentic AI reliability is the critical unsolved frontier. The Nemotron 3 Ultra test is emblematic—models can generate coherent multi-file code and stay on-task for hours, yet fundamentally cannot verify their own output. The model's blindness to its own failure (declaring victory on an empty screen) is a systemic issue across frontier models, not vendor-specific.

Escalation: Local AI is transitioning from enthusiast niche to production infrastructure. Multiple converging signals—NVIDIA's free hosted access, Framework's local inference hardware, dedicated summits, Hermes' rapid feature expansion—indicate Local AI is building real tooling. However, @sudoingX's critique that every "local AI for everyone" tool still requires engineering expertise highlights the remaining gap.

De-escalation: The token maximization era is cooling. ROT (Return on Tokens) formalizes what practitioners have been feeling: throwing more tokens at problems is yielding diminishing returns. Smart routing and token efficiency are replacing brute-force approaches, mirroring Coinbase's reported cost reductions.

What to watch: Grok's teased major upgrade; whether NVIDIA addresses Nemotron's visual verification gap; Local AI Summit outcomes on Thursday; Hermes Agent's continued feature velocity; and whether "proof of individual" identity systems gain traction as AI-generated content proliferates.

Tweet Feed

NVIDIA Nemotron 3 Ultra & Model Testing

@sudoingX · 2026-06-29T18:57

watch this anon. i gave NVIDIA's biggest model ever a single task. 100 minutes and 440,000 tokens later, it had rendered nothing. not one important thing on the screen. this is Nemotron 3 Ultra. 550 billion parameters, a hybrid Mamba Transformer MoE, the largest model NVIDIA has ever shipped, and they built it specifically for long-running agentic coding. so i handed it exactly that: build a 3D scene from a spec, multiple files, iterate until the tests pass. the same task a frontier model one shotted in minutes. i genuinely wanted to be impressed. it ran for an hour and forty. burned through 440,000 tokens. wrote every file, passed its own tests, and proudly printed "task complete."the browser was blank. the 3D scene never rendered. not once. and the long horizon agentic behavior was genuinely good. it stayed on task the whole hour and forty, wrote real multi-file code, drove its own tools without derailing. it just couldn't turn any of that into something that actually runs. here's the part that gets me. it's a text model, it cannot see its own output. so it sat there looping on a broken vision tool, trying to "look" at the page, hitting error after error, never once reasoning its way out. it declared victory on an empty screen because it had no way to know the screen was empty. → tweet link

@sudoingX · 2026-06-29T12:58

this is Nemotron 3 Ultra. 550B total, 55B active, the biggest MoE nvidia has ever shipped in house. and it's not big just to be big, it's a hybrid mamba transformer MoE built specifically for long running agentic workflows, the kind of coding and research runs that go for hours and fall apart the second a model loses the thread, which happens to be exactly the kind of work i put every model through, so this one lands right in my wheelhouse. i've been hearing about it for weeks, that it punches way above its weight, that it trades blows with frontier models people pay a fortune to touch, and i'm not the type to take nvidia's word for it or anyone else's, so tonight i'm running it through its paces myself. it's way too big to run local, 550B needs something like 4 DGX Sparks just to hold the weights, so i'm hitting it through nvidia's NIM and driving it with Hermes Agent, the same harness i run every model through, which means this isn't a clean little demo, it's the 550B doing real agentic work under my conditions on my prompts. → tweet link

@sudoingX · 2026-06-29T12:35

have to give nvidia their credit. they set me up with an NGC account, free access to their whole model lineup, including Nemotron 3 Ultra, their 550B flagship. open weights, open data, and now free hosted access so builders can actually run them. while the big labs gatekeep harder every quarter, nvidia quietly became one of the most open players in AI. i love the direction they're moving. respect where it's due. → tweet link

Local AI Movement

@TheAhmadOsman · 2026-06-29T16:20

RT @TheAhmadOsman: MASSIVE NEWS. Teamed up with NVIDIA to make Local AI The Default → tweet link

@TheAhmadOsman · 2026-06-29T07:55

People are massively sleeping on this guy in the Local AI space btw → tweet link

@TheAhmadOsman · 2026-06-28T20:02

My mission since 2023 has been to teach people and prepare them running their own AI. June 2026 marks the most important chapter since I started on this mission. Watch us make Opensource and Local AI the default → tweet link

@TheAhmadOsman · 2026-06-28T21:52

The opposite of Anthropic is … → tweet link

@TheAhmadOsman · 2026-06-28T23:50

If it can be automated with a Cli or accessed through an API Endpoint, I hand it over an agent with GLM 5.2 to do it for me → tweet link

@TheAhmadOsman · 2026-06-29T02:57

We actually kickstart things early tomorrow with 2 workshops on Local AI. See you guys tomorrow → tweet link

@sudoingX · 2026-06-28T19:22

i see it constantly now, everyone "making local AI the default," launching platforms to bring it to the masses. and i've yet to see a single one that actually solved the infra for real users. here's the irony nobody says out loud: almost every "local AI for everyone" tool still needs you to be a computer expert just to stand it up. you have to already know the thing it claims to remove the need for. that's not solving infra, that's a dev kit with better marketing. the actual problem is making local AI work for someone who isn't an engineer, is still wide open. i see it clearly. → tweet link

@sudoingX · 2026-06-28T19:37

almost 3am anon, and before i crash, a reminder: if you're getting into local AI or agentic developments, hermes agent is the leanest door in. stands up in seconds, no onboarding maze. you just run it, and you're in. start with the thing that gets out of your way. night. → tweet link

@sudoingX · 2026-06-28T21:20

welcome herald. hermes agent everywhere. → tweet link

Hermes Agent & OpenCode

@Teknium · 2026-06-28T19:46

Should we add a wake word voice chat starter into the desktop app? ^_^ → tweet link

@Teknium · 2026-06-28T22:26

Tip of the day, the CLI is fully skinnable, so you can customize your Hermes Agent look and feel! → tweet link

@Teknium · 2026-06-29T05:47

Nice video by @WesRoth on Hermes' Stripe integrations! Check it out: → tweet link

@Teknium · 2026-06-28T22:28

RT @ElevenLabsDevs: How to talk to your Hermes agent and hear it talk back. → tweet link

@Teknium · 2026-06-29T03:44

Someone needs to find the right workflow to get Hermes Agent to finish the Berserk 1997 anime.. → tweet link

@thdxr · 2026-06-29T18:09

in OpenCode v2 there is a new subagent/shell management ui. you can see what's running and easily background or kill it → tweet link

@thdxr · 2026-06-28T22:58

i told opencode to use my browser to sign up for telnyx and give itself a number. they had agent instructions for sign up! it even had this crazy bot challenge. the rest of the onboarding was tailored for my agent and it got a number and got everything setup → tweet link

@thdxr · 2026-06-29T17:52

RT @opencode: OpenCode Go users in New Zealand used 40.9B tokens last week, or, 1,734 tokens per sheep → tweet link

Return on Tokens (ROT) & Token Economics

@LinusEkenstam · 2026-06-29T00:25

ROT — Return on Tokens. We all knew we would end up here at some point. Tokenmaxxing was a dumb idea from the start. It's just that when everyone is drinking from the coolaid it's hard to resist. But for some of us it's been clear as day for years that first we will see unhinged growth no matter the cost. Because nobody will want to be left behind. Leaderboard with top token spenders as if it was a direct correlation to output and value. When in reality it's a spend leaderboard and nothing else. […] Being able to route well is a key lever to good RUT. While dropping usage is the simplest, routing might be the hardest. Because of its near infinite possible outcomes. → tweet link

@LinusEkenstam · 2026-06-28T23:58

Everyone had access to Fable. Almost nobody took advantage of it. During that short window, I somehow did. It made over $43,000 for me. 🤯 I can see how it's a national security risk → tweet link

AI Hardware & Infrastructure

@FrameworkPuter · 2026-06-28T20:42

MTP is another great way to crank up inference speeds on the Ryzen AI Max in Framework Desktop, especially for code generation. @dcapitella has some great numbers to show with this. 6.3 tok/s -> 13.7 tok/s with Qwen3.6-27B-Q8. → tweet link

@gospaceport · 2026-06-28T22:29

Great Question @NakeZast and here are 3 vids on the rundown on how I built my Home Datacenter's Power and Cooling. It works fantastic. → tweet link

@TrungTPhan · 2026-06-29T01:39

RT @bearlyai: Bloomberg chart showing amount of RAM needed for AI data centres. Integrated server rack of 72 Nvidia Blackwell chips = same… → tweet link

@TrungTPhan · 2026-06-29T17:43

RT @bearlyai: Nvidia and Eli Lilly are building a $1B lab for AI drug discovery in SF. Lilly CEO Dave Ricks talks about their data advantage… → tweet link

@gospaceport · 2026-06-29T02:32

Grok getting a major upgrade → tweet link

Developer Tools & Open Source

@kunchenguid · 2026-06-29T18:21

been shipping a ton of improvements to firstmate. a fun one is that i can now directly talk to my local firstmate from X. ahoy @myfirstmate - give us a concise recap of what we shipped to firstmate repo over the last week? this is your first public appearance so make it good → tweet link

@kunchenguid · 2026-06-28T19:41

here's the final results from the poll for "the best 3rd subscription". cursor is overwhelmingly the most popular recommendation. "other" is a combination of for GLM, deepseek, ollama cloud, grok, and a long tail of choices → tweet link

@alexinexxx · 2026-06-29T14:16

espanso >dropover >raycast >zen browser >ghostty. what else should i install on my macbook? → tweet link

@victormustar · 2026-06-28T21:51

RT @fal: We just open-sourced 3DREAL! A new render-to-real IC-LoRA for LTX-2.3 from fal: turn any 3D / CG / game render into high 3d ren… → tweet link

@jezell · 2026-06-28T23:03

RT @tom_doerr: Modular GraphRAG implementation in Rust with WebGPU acceleration support. → tweet link

@victormustar · 2026-06-29T06:07

RT @_alejandroao: introducing tau τ — an educational agent harness that teaches you how to build agent harnesses. i will be publishing tuto… → tweet link

@badlogicgames · 2026-06-28T20:25

omg it's basically pi, but more minimal for educational purposes and in python. @marlene_zw is gonna love this. great educational resource, alejandro! → tweet link

@levelsio · 2026-06-29T15:46

✅ Added to [his site]. You can now browse 102,913 Winamp skins from The Internet Archive and load them directly! → tweet link

@iamdevloper · 2026-06-28T20:04

  1. Learn Rust → tweet link

@alexinexxx · 2026-06-28T19:50

RT @yacineMTB: I need to ignore everything and focus on studying GPU research papers published by Tero Karras → tweet link

AI Agents & Web Accessibility

@steipete · 2026-06-29T02:06

RT @cherry_mx_reds: AI agents may end up doing more for web accessibility than accessibility laws. Agents rely on accessibility APIs to use… → tweet link

@LinusEkenstam · 2026-06-29T11:17

Just like that we just super charged the dead internet theory even further. I can't wait for proof of individual to create accounts on social media and full transparency if you operate multiple accounts. → tweet link

@steipete · 2026-06-29T00:46

RT @HamelHusain: Codex desktop is miles ahead of Claude for remote access via mobile or another computer. Codex allows you to see *all… → tweet link

AI Economics & Data

@cooltechtipz · 2026-06-29T07:06

PwC estimates AI could add $15.7 trillion to the global economy by 2030. → tweet link

@cooltechtipz · 2026-06-29T14:57

AI Data Center Networking. → tweet link

@cooltechtipz · 2026-06-29T10:10

AI creates opportunities. How much each country gains depends on how well it prepares. The IMF says countries investing in AI, computing, and workforce skills are likely to benefit the most. → tweet link

@FinansowyUmysl · 2026-06-28T20:16

RT @FinansowyUmysl: Bardzo ciekawy case study jak Coinbase zmniejszyło rachunki za AI, prawie o połowę, nie zmniejszając zużycia tokenów. → tweet link

Agent-First Video & GenUI

@LinusEkenstam · 2026-06-29T11:10

RT @neiltak: Building an agent-first video editor with motion baked in. Yes, it has shaders, plugins, multi-comp canvas, shareable componen… → tweet link

@ASalvadorini · 2026-06-29T09:21

RT @ulusoyapps: Another before & after @FlutterDev GenUI demo 🤓. This time Generative UI #GenUi replaces model's writing practice feedback… → tweet link

Miscellaneous Tech

@levelsio · 2026-06-29T10:42

Windows is honestly annoying to use, every time I open my PC to play Flight Simulator 2024 I get some dumb popups like this, either about selling me OneDrive to backup my drive or now upgrading to paid Microsoft 365 (I don't even use office apps on this PC, it's for gaming!) And it's not just one button to skip or decline, it's multiple, plus the dark pattern of this [ Upgrade ] CTA which looks like [ Continue ] so I almost did upgrade accidentally. Considering I already paid $130 for Windows 11 Pro just to use it, it's crazy to keep getting upsold inside your OS. MacOS never does this and also I never had to pay for MacOS → tweet link

@levelsio · 2026-06-29T12:00

What people suspect is indeed true. Negative content performs much better. About 1.5x better than positive content. But curious content (like hacky projects) comes in second, which is great news for me! → tweet link

@Ex0byt · 2026-06-29T13:36

new JARVIS unlock → tweet link

@gdb · 2026-06-28T22:09

ChatGPT for helping in daily life in Bengaluru → tweet link

@sudoingX · 2026-06-29T10:40

the proxy infra my platform runs on is maxed out and i need a new server to keep it stable. → tweet link

@nummanali · 2026-06-28T20:08

I have an exceptional AI Native Full Stack Engineer who's been trained under me. If anyone is looking for a new core team member, DMs are open. Requirements: Remote, £70K+, AI First company, A team that cares, Mission driven co → tweet link

@alexocheema · 2026-06-29T01:44

Anyone at AIE badge pickup (Moscone) want an exclusive sneak peak of [project]? I have it on my laptop here. → tweet link