← Tech / AI / IT Monitor Index Tech / AI Generated 2026-08-09 19:30 UTC

Tech / AI / IT Monitor

August 09, 2026 · Based on tweets from the last 24 hours · 126 tweets analyzed · model: ollama-cloud/glm-5.2:cloud

Executive Summary

The last 24 hours saw significant focus on local LLM inference and hardware optimization, particularly around running DeepSeek v4 Flash on Apple Silicon and NVIDIA DGX Sparks. ByteDance's Seedance 2.5 video model also launched, marking a massive leap in AI-generated video realism and character consistency. Meanwhile, developer tooling evolved with Anthropic and OpenAI introducing new workflows, such as Claude Code agent-to-agent messaging and OpenAI web search integrations, showing a clear trend toward complex, autonomous coding agents.

Key Events

Analysis

A prominent pattern across the timeline is the growing push toward "owning your thinking"—developers are actively moving away from rented, frontier API models in favor of local, open-weight models (like DeepSeek v4 Flash and Qwen 3.x) running on high-end personal hardware (Apple Silicon, DGX Sparks). This is driven by a desire for privacy, predictable costs, and model persistence without sudden deprecations. Another trend is the maturation of agentic workflows; the conversation is shifting from simple code generation to building complex verification pipelines, agent-to-agent communication, and UI/UX integrations (like HUD modes). What to watch next: The impending release of Qwen 3.8 27B and how dense models perform on unified-memory architectures like the DGX Spark.

Tweet Feed

AI Models & Hardware

@ivanfioravanti · 2026-08-09T18:57

I'm new and a super noob on AMD AI space, but... how is it possible that Kernel 7 is not officially supported by drivers? 🤔 https://t.co/j7Dd9q34hO → tweet link

@ivanfioravanti · 2026-08-09T18:36

I wonder what is the speed of DeepSeek v4 Flash-0731 on 3 x DGX Sparks 🤔 We'll discover it soon 😎 https://t.co/cKFusj8zQU → tweet link

@ivanfioravanti · 2026-08-09T18:39

RT @antirez: DwarfStar with DFlash speculative decoding now can do the DeepSeek v4 Flash inference much faster both when used with Metal… → tweet link

@ivanfioravanti · 2026-08-09T15:40

Hybrid AI Experiments with MiniMax H3. 2K generation online costs $0.13/sec... Just some initial experiments. API docs here: https://t.co/0sY5IgweEr → tweet link

@Ex0byt · 2026-08-09T13:57

woke up to 2 new MoE surrogate breakthroughs with algebraic proofs, wired (sglang) and running against DSv4-Flash-0731 on a donated rig... what a time to be alive. https://t.co/3l3vSHTHHU → tweet link

@ivanfioravanti · 2026-08-09T13:44

Quick context benchmark test up to 64K using DwarfStar Apple Silicon decode-optimized branch versus baseline running on M3 Ultra with DeepSeek v4 Flash 0371 mxfp4 and temperature 1... 64k pp 484 tg 35 t/s → tweet link

@sama · 2026-08-09T14:09

RT @gdb: GPT-4 finished training four years ago today. → tweet link

@ivanfioravanti · 2026-08-08T22:44

Ok! My new ds4 fast branch is ready! There was a bug in RoPE finally solved! Ready for the prime time! ~44 tok/s on M3 Ultra mxfp4 ~46 tok/s on M3 Ultra Q2 ~ 45 tok/s on M5 Max Q2 → tweet link

@ivanfioravanti · 2026-08-08T22:23

Next week we'll get Grok 4.6 and Qwen 3.8 27B! → tweet link

@TheAhmadOsman · 2026-08-08T21:00

Dense models like Qwen 3.8 27B are a TERRIBLE experience on unified-memory systems like DGX Spark btw. DGX Sparks are better suited to MoE models with relatively few active parameters per token → tweet link

@RayFernando1337 · 2026-08-08T19:42

RT @TechMDAI: Built my first hybrid vision model around DeepSeek V4: DeepSeek V4 Flash Hybrid Vision. → tweet link

AI Video & Image Generation

@levelsio · 2026-08-08T22:53

As always everyone is blind staring at the progress of LLMs for coding and chat. But meanwhile the new SOTA video model Seedance 2.5 has been slowly rolling out... I'd say it's the first video model that's now at the level of image models with the level of character likeness... → tweet link

@FinansowyUmysl · 2026-08-09T16:50

Lubię testować nowinki AI - a jeszcze bardziej takie, które realnie przyspieszają tworzenie. Seedance 2.5 robi dokładnie to. Jednym promptem możesz wygenerować profesjonalnie wyglądające video... → tweet link

@levelsio · 2026-08-08T23:11

I present you: iPhone 18 Pro Fold ❤️ ... I uploaded a clip of my voice, and a few pics, and this came out, it's getting close to passable! Live on Photo AI now 😊 → tweet link

@levelsio · 2026-08-08T23:22

I tred pumping this blog post into it... we're getting pretty close to AI-gen'd YouTubers. This took 5 minutes to generate! → tweet link

Developer Tools, Frameworks, & Agents

@swyx · 2026-08-09T17:32

occasional reminder to DELETE your skills. when you are bombarded constantly by "this skill changed my life!!" on the timeline, you will pile up stuff that at best just eats context, and at worst interacts with other skills nastily... → tweet link

@TrungTPhan · 2026-08-09T17:22

RT @bearlyai: [NEW] Bearly AI users can now create Routines: describe a task in plain language (eg. “a morning brief of AI news”, “track s… → tweet link

@jezell · 2026-08-09T17:51

three_flocker running 500+ three.js samples on MacOS via WASM + Dawn and Flocker. Host app is standard Flutter app, three_flocker is a dart port of the undisputed champion of 3d on the web @threejs that compiles to WASM and renders with WebGPU across web, mobile, and desktop. → tweet link

@jezell · 2026-08-09T17:04

FlockerPad running on iOS, compiling a flutter app by launching dart2wasm as a plan9 child process and then using the plan9 view device to embed launch the wasm output and render in a child view in the parent app. → tweet link

@gdb · 2026-08-09T18:13

Codex for saving money by reading the fine print: → tweet link

@jxnlco · 2026-08-09T15:59

codex does not have a /loop command, but if you just send that string, it knows what to do → tweet link

@ivanfioravanti · 2026-08-09T14:23

Let them (agents) cook! I need to use all my Codex sub by the end of the day! https://t.co/jRkuxb9KMH → tweet link

@Teknium · 2026-08-09T08:14

The Hermes Agent CLI will now display the name the session gets in the right side of the context bar area :) https://t.co/lTUQy7cmCB → tweet link

@gdb · 2026-08-09T08:11

ChatGPT Finance for helping you save money: → tweet link

@swyx · 2026-08-09T07:07

i shipped first set of llm-as-judge evals for the kill my saas competition tonight. people can run this to check if their solutions at least pass the sniff test. https://t.co/2eYVts1Yhp → tweet link

@swyx · 2026-08-09T05:31

i still think @AnthropicAI ultracode is one of the most important coding mode innovations ever invented. if you havent understood the potential of dynamic workflows you should try to. → tweet link

@Teknium · 2026-08-09T02:33

Tomorrow we save you 60% tokens on browser use tasks :] → tweet link

@Teknium · 2026-08-09T00:30

Can confirm I have 2x sparks connected with one cable and get around 40tok/s without dspark on deepseek v4 flash 0731 abliterated. Completely uncensored, completely private inference. → tweet link

@Teknium · 2026-08-09T00:25

RT @jonkomet: Hermes Browser Extension🪽 just got a theme studio. this one brings real theming, you can now pull any VS Code look into your… → tweet link

@thdxr · 2026-08-08T23:59

RT @Neriousy: openai websearch 🤝 @opencode. users with chatgpt subscriptions can now use the openai websearch inside opencode → tweet link

@badlogicgames · 2026-08-08T20:37

RT @alonwo: Claude Code now lets its sessions message each other (agent-to-agent, via /list-agents). I built pi-claude-link so Pi sessions… → tweet link

@jezell · 2026-08-08T20:49

RT @thsottiaux: That's right, GPT-5.6 Sol is awesome and can be used pretty much anywhere, including in the CC harness. → tweet link

Local AI & Workflows

@sudoingX · 2026-08-09T04:10

if you ever feel behind, just remember some people are still paying $200/month for a coding agent that a used 3090 and an open model do for the price of electricity. → tweet link

@sudoingX · 2026-08-09T04:16

most companies i've worked with have never once benched their actual workload against a local model. they pay frontier api prices for tasks a 27b would clear on a single gpu card. → tweet link

@sudoingX · 2026-08-08T20:01

the ai bubble is smaller than it looks... owning your thinking starts with owning metal. weights on your own disk run identical in ten years, nobody can revoke, reprice, or rewrite them. → tweet link

@sudoingX · 2026-08-08T19:11

we've reached the stage where the harness matters more than the base model but the intelligence is leaking out of the model and into the loop around it. in two years nobody will ask what model you use. → tweet link

@ivanfioravanti · 2026-08-08T19:11

I don't know who's the culprit here, but running out of memory on an M3 Ultra 512GB is nearly impossible with no running AI workload. 🧐 I smell a memory leak somewhere. → tweet link

@jezell · 2026-08-08T19:22

It's really amazing what LLMs can do if you focus your instructions and effort more on the verification and vetting process. You need to build data driven pipelines with objective, repeatable, and measurable gates. → tweet link

@MatejKnopp · 2026-08-08T20:17

When I vibe-code a 50 line PR I have claude rewrite it 4 times until I'm happy. But 3000-line vibe-coded feature PR? Seems fine to me, ship it. Does this really not bother anyone? → tweet link

@ivanfioravanti · 2026-08-09T09:40

Personal suggestion for code/PR review. Ask the model to give you a detailed html report on it and when you review/fix things, point the agents to each html section and ask it to update them once done. → tweet link