Executive Summary
The past 24 hours saw significant advancements in local AI hardware capabilities and agentic coding frameworks. Open-source models like DeepSeek V4 Flash and Ling 3.0 Flash are demonstrating remarkable efficiency on consumer-grade hardware like the NVIDIA DGX Spark, enabling 100B+ parameter models to run on local boxes. Agent infrastructure is rapidly maturing, with Prime Agent achieving near-human-expert baselines on ARC-AGI-3 and Nous Research's Hermes Agent expanding its desktop and document-handling capabilities. Additionally, novel development environments like FlockerPad are emerging, leveraging WebGPU and Flutter to rethink cross-platform IDE architectures.
Key Events
- Prime Agent scores 95.5% on ARC-AGI-3, surpassing the human-expert baseline using a fully open-source pi-based harness. → link
- NVIDIA DGX Spark users report running DeepSeek V4 Flash (284B) locally at 16.5 tok/s and Ling 3.0 Flash (124B) at 40 tok/s. → link
- Nous Research's Hermes Agent gains desktop browser capabilities and universal document-to-Markdown conversion powered by Firecrawl's open-source Rust library, anydoc. → link
- FlockerPad is announced, utilizing Plan9, Flutter, Wasm, TreeSitter, Lapce, and WebGPU to build a cross-platform code editor without traditional JS hacks. → link
- The White House exempts open models from its new framework to test frontier AI capabilities before release. → link
Analysis
There is a clear pattern of decentralization and localization of AI compute, with developers increasingly running 100B+ parameter models locally on unified memory architectures. Simultaneously, the "harness" layer—the orchestration tools that manage agents—is becoming a critical battleground, as seen with the rapid feature expansion of Hermes Agent, Prime Agent, and OpenCode v2. The industry is also coalescing around practical agent standards and routing mechanisms, balancing high-cost frontier models with efficient open-weights for specific sub-tasks. Watch for further convergence on agentic memory graphs, self-improving RL harnesses, and the maturation of WebGPU-based local execution environments.
Tweet Feed
AI Models & Hardware Benchmarks
@sudoingX · 2026-08-06T03:55
i'm looking at what this one dgx spark box actually runs. 128gb unified. sits beside my monitor.
every model on it, biggest to smallest, with the quant and what it did smoke the shit out of:
DeepSeek V4 Flash, 284B / 13B active. IQ3_XXS 3-bit, fully on the gpu, zero offload. 16.5 tok/s. a 284B model resident on a desk box, usable.
StepFun 3.7 Flash, 198B / 11B active. ran the Q4_K_S gguf at 262k context, 25 tok/s single-stream. their readme claims 279, i measured 25, off by 11x. the NVFP4 build (121gb) won't even fit the box. the honest edge of what 128gb holds.
Qwen 3.5 122B, 122B / 10B active. NVFP4. 28 tok/s base, 35 with MTP. loaded first try, ~30gb of kv cache to spare, 18x concurrency at 128k. → tweet link
@sudoingX · 2026-08-06T13:54
a 124B model doing 40 tok/s on a box beside my monitor. i benched Ling 3.0 flash, @AntLingAGI's newest open model, across the community quant ladder on one dgx spark, and here's what it actually does.
decode, single stream:
Q5_K_M, 40.2 tok/s, the sweet spot, fastest AND near-lossless Q4_K_M, 38.2 tok/s, the speed pick, smallest footprint Q6_K, 32.0 tok/s, max quality for about 16% off the top
the whole ladder sits in a tight 32 to 40 band. the sweet spot is 2.4x deepseek v4 flash on the same box, 40 vs 16.5 tok/s, and even max quality Q6 is nearly 2x. so few params fire per token that the quant barely moves it. → tweet link
@sudoingX · 2026-08-06T09:20
pulling Ling 3.0 flash onto my dgx spark, and hear this part almost nobody tells you about running the newest models: on day one you often can't even load them.
Ling isn't a normal transformer. @AntLingAGI built it on a native hybrid linear architecture, five Kimi-Delta-Attention layers for every one MLA layer, 124B total but only 5.1B active per token. → tweet link
@louszbd · 2026-08-06T07:47
Looks like Muse Spark 1.2 made strong gains on long-horizon knowledge work. Impressive progress from the Meta team! https://t.co/qmN27fF8Kk → tweet link
@thdxr · 2026-08-06T14:49
on the upcoming deepseek price increase
we've been able to reproduce their current prices even on rented GPUs
so this likely isn't because they're "losing money" it's traffic shaping because they are overloaded → tweet link
Agent Frameworks & Harnesses
@Teknium · 2026-08-06T00:11
Hermes Agent can now read anything you throw at it. PDF, Word, PowerPoint, Excel, OpenDocument, RTF, EPUB — all auto-converted to clean Markdown the moment the agent reads the file, all locally!
Powered by @firecrawl_dev's new open-source anydoc (Rust). Zero setup, installs itself on first use. → tweet link
@Teknium · 2026-08-05T21:06
Discover how your Hermes Agent grew with you, by exploring /journey's memory graph showing a timeline of your memories and its self improvement loop's skill development! → tweet link
@ivanfioravanti · 2026-08-06T12:16
So far so amazing good with Prime Agent! "Hint: Prime Agent self-improves by refining skills, memories, prompts, and subagents."
This is true! And you see this iterating on the same project and tasks multiple times.
Great job @PrimeIntellect → tweet link
@badlogicgames · 2026-08-05T19:50
RT @PrimeIntellect: Prime Agent was designed as a coding agent, but also can be used for general agentic tasks in any domain.
Prime Agent… [continues] → tweet link
@thdxr · 2026-08-06T05:35
opencode v2 knows its own api
i pasted a slack convo from a v2 tester with a bunch of issues they mentioned and it made a session for each problem with an initial prompt
opened them all up in tabs and now i can work through them https://t.co/nUFWlzSiuz → tweet link
@kunchenguid · 2026-08-06T18:08
agent plugins - this is a good and much needed standard
but... it's only truly useful if we can get anthropic to also follow. otherwise we still have to publish separate plugins
do you think Anthropic will converge? → tweet link
@swyx · 2026-08-06T06:33
a very primitive form of the near term multiagent agi future is setting up one thread to ping back once its done so you create an implicit kanban/waterfall graph of dependent threads but each preserving their own work and agents → tweet link
Developer Tools & Platforms
@jezell · 2026-08-06T17:47
Plan9 + Flutter + Wasm + TreeSitter + Lapce + WebGPU = FlockerPad. Flutter app is run as a Plan9 child process. Dart2wasm.wasm is compiling the code. Same code runs on web, mobile, desktop. Virtual file system in the browser with support for dart:io. https://t.co/DiGo0tqCaw → tweet link
@jezell · 2026-08-06T06:10
I'll tell you one thing. FlockerPad is gonna be Dart. It ain't gonna be a CodeMirror wrapper like DartPad. If you can't even build a decent code editor on your GUI framework, you should not be trusted. TreeSitter running as a plan9 device in a background process, Lapce flutter port doing the edits. The flocker device model would actually be awesome for running LSPs across platforms... → tweet link
@badlogicgames · 2026-08-06T15:37
RT @mitsuhiko: People of pi: a significant new release, 0.84.0 is out. In settings you can now turn on fullscreen (alt screen) mode. Pi now… → tweet link
@levelsio · 2026-08-06T15:47
I think terminal based coding agents on the server are way more powerful than LLM apps because they can do actual stuff on your server like optimizing your Nginx config, speed up your SQLite db, or fix Ubuntu stuff like automatic upgrades → tweet link
@victormustar · 2026-08-06T14:35
RT @rough__sea: Introducing celld: a self-hosted, distributed Durable Objects and Workers implementation
- celld = V8 + S3 + SQLite + LTX+… → tweet link
AI Ecosystem & Industry Trends
@LinusEkenstam · 2026-08-05T22:34
Just before bed time. Let me sleep. plz
95.5% on ARC-AGI-3 🤯
(’huge if true”) → tweet link
@jsuarez · 2026-08-06T04:44
RT @MTSlive: SITUATION UPDATE: The White House has exempted open models from its new framework to test frontier AI capabilities before rele… → tweet link
@LinusEkenstam · 2026-08-06T07:15
I’m convinced that we’re building more and more dependencies on LLM’s to the point at which the smallest degradation will feel like the end of the world.
It’s also insane how a small weight somewhere has turned the model into a lazy, degraded co-worker instead of the A-player it was just a short while ago. → tweet link
@MilksandMatcha · 2026-08-05T20:13
Not every agent task needs a frontier model.
@Thom_Wolf (Co-founder of Hugging Face) on routing work to smaller open models for repeatable sub-tasks, evaluation, and cost control, matching compute to the value of the work. → tweet link
@gdb · 2026-08-05T20:06
full house for the team’s talk at Black Hat on the OpenAI-Hugging Face Incident https://t.co/vf52ssLZZc → tweet link
@victormustar · 2026-08-06T13:49
RT @allen_ai: We're expanding our partnership with @huggingface to accelerate open science.
Our storage on the Hub is roughly tripling to… → tweet link
@levelsio · 2026-08-06T16:14
Meta is doing heavy heavy heavy scraping on all my sites this week too (and seemingly everyone else's sites now)
So much so that I got load average alerts for it today on one VPS
The fact that they're hitting url2og (my own screenshot service for all my sites) might they're scraping not just text but images too, so maybe they're working on an image, video or just world model → tweet link