← Tech / AI / IT Monitor Index Tech / AI Generated 2026-08-08 19:30 UTC

Tech / AI / IT Monitor

August 08, 2026 · Based on tweets from the last 24 hours · 162 tweets analyzed · model: ollama-cloud/glm-5.2:cloud

Executive Summary

OpenAI teased its next major model, "Astra," which shows significant advancements in agentic coding and cybersecurity, currently pending safety reviews before broad release. In the open-weights and local AI space, major efficiency breakthroughs were achieved as AntLing's 124B Ling 3.0 Flash was successfully optimized to run on a single DGX Spark at ~38.7 tok/s, while Ollama rolled out DeepSeek-V4-Flash-0731 on its cloud with over 200 tps. Developer tooling and agent ecosystems also expanded rapidly, highlighted by the evolution of Hermes Agent, new Apple Silicon optimizations (DwarfStar), and the addition of lazy tiered JIT to Wasmtime for 219x faster WASM startup. Additionally, Qwen 3.8 27B weights were teased for next week, promising frontier-level intelligence running locally on consumer MacBooks.

Key Events

Analysis

The past 24 hours highlight a strong push towards high-performance local AI and agentic self-sufficiency. Innovations in quantization and kernel optimizations (like AntLing's int4 path and DwarfStar's M5 updates) are enabling 100B+ parameter models to run efficiently on prosumer hardware (DGX Spark, Framework Desktops, MacBooks). There is a growing trend of developers bypassing cloud APIs for privacy, cost, and latency reasons, instead using local agents like Hermes to write software and manage system tasks autonomously. On the proprietary side, OpenAI's tease of "Astra" signals a continued focus on cyber capabilities and agentic coding, while open-weight models (DeepSeek, Qwen) are aggressively challenging price-performance ratios to eliminate "agent spending problems." Watch for the Qwen 3.8 release next week, as well as further safety updates from OpenAI regarding Astra's deployment.

Tweet Feed

AI Models & Releases

@gdb · 2026-08-08T03:15

luna is such a special model, incredible price performance → link

@gdb · 2026-08-07T19:11

Evaluations of our next major model, Astra, indicate significant capability advancements in agentic coding and cybersecurity.

Team is doing the safety and security work to make Astra broadly available, and get its advanced cyber capabilities into the hands of defenders: → link

@sama · 2026-08-07T22:54

astra is a powerful model and we are working to make it generally available.

we do not think it is a good strategy to keep powerful models to a chosen few.

given its cyber capabilities, we need a little big longer to do do this safely. but hopefully not too long! → link

@ollama · 2026-08-08T06:24

Currently serving 200 tps+ output speed for DeepSeek-V4-Flash on Ollama's cloud with zero data retention (ZDR).

Have an amazing weekend 🫡 → link

@ollama · 2026-08-07T19:54

DeepSeek-V4-Flash-0731 is now fully rolled out as the new default for deepseek-v4-flash on Ollama's cloud.

This model combines speed, efficiency, and frontier-level performance.

Fast: 120+ output tps on Ollama's cloud Private: zero data retention hosting in US & Europe Efficient: generous usage on Ollama's Pro and Max plans for multiple long-running, uninterrupted sessions with your favorite coding harnesses. → link

@alexocheema · 2026-08-07T20:04

Qwen 3.8 27B weights are coming out next week.

It's a big deal.

We are going to have close to GLM-5.2 level intelligence running locally on a 36GB MacBook ($4k) at ~25 tok/sec.

Watch this so you know what's coming. Great convo w/ @jason @Lons @eisokant @Ape → link

@Teknium · 2026-08-08T10:32

With where deepseek flash is on price performance i feel pretty strongly that the ai agent spending problems will go away quite quickly for those that want it to → link

Local AI & Hardware

@sudoingX · 2026-08-08T15:51

if you own a dgx spark and want to run @AntLingAGI's official ling 3.0 flash quants, i open sourced the whole thing.

the official int4 through their vllm fork is now the fastest way to run this model on a single spark, 38.7 tok/s, past the community gguf everyone defaults to. stock vllm will silently give you garbage, the repo explains why and what to run instead.

clone it, serve it, you're generating in minutes: https://t.co/suQQbcKA15 → link

@sudoingX · 2026-08-08T15:15

i need to correct something i posted yesterday. i told you antling's official quants don't run on a single dgx spark.

tonight the @AntLingAGI official int4 is the fastest way to run ling on one, and the same box just built me a working snake game with this model.

let me back up, because this one genuinely got me. ling is a 124B model, and yesterday every official quant AntLing shipped choked on one spark, fp4 wanted two boxes, fp8 wouldn't fit, the int4 loaded and then died inside a gpu kernel. i was ready to call it a two box model and move on.

so i went into their own code, and the speed was already there, just unwired, their vllm fork, their KDA kernels, their MTP draft layer, all shipped, none of it bolted together for this box. stock vllm runs it on the wrong attention math and hands you fluent garbage that reads fine until it doesn't, the naive config leaves half the speed on the floor, the shards freeze on every other cold start.

i faced three walls in one night, and not one of them was the model's fault.

and when wired right, the official int4 does 38.7 tok/s on one spark, past the community gguf everyone actually runs at 35.2, and it serves the model's full 256K context window on the same box.

and watch. i pointed hermes agent at ling 3.0 flash on this box and gave it one prompt, and it wrote a complete snake game in one shot, verified its own file, and opened it to play. a 124B model, on a box beside my monitor, building working software on its own.

a model that wouldn't load yesterday is the fastest path on the hardware today, open weights, the whole recipe public. best kind of week there is, and it only gets better from here.

recipe, scripts, every wall documented below 👇 → link

@ivanfioravanti · 2026-08-08T15:55

BREAKING: DwarfStar optimizations now on M5 kernel too! I focused on decoding and Q2 quantization here:

40 tok/s --> 45 tok/s ! 🔥

I recorded videos (slowing down decoding speed) with baseline and optimizations. I fixed tool calls loop too here! M3 Ultra fix still in progress.

Branch is here: https://t.co/4JC7sKLSqt

Now closing fix for M3 Ultra, cleaning code, another full round of tests and if everything is ok I'll open PR for super @antirez to review 🚀 → link

@sudoingX · 2026-08-08T09:36

i got HunyuanImage 3.0 running on a Framework Desktop, 80B parameters, 83GB of int8 weights in 128gb unified memory on a Ryzen AI Max+ 395.

first thing i asked for was a black monolith. these four are from about thirty attempts. shot 1 is the only one with an object in it, and it took ten tries to get the slab to stay a slab.

the interesting part is where it sits on the curve. this is what 2026 gives you at home on hardware you can buy off a shelf, and the models are still getting smarter at the same size while the memory gets wider.

local image generation in 2026 looks like this. → link

@TheAhmadOsman · 2026-08-08T17:53

By the way

RTX PRO 6000 = 1.8TB/s

DGX Spark = 273GB/s

You should aim for higher bandwidth if agents and agentic swarms are your goal with Local AI - it matters a lot → link

@ivanfioravanti · 2026-08-08T07:23

Playing with Wan Animate 2 by @Alibaba_Wan team and ComfyUI on DGX Spark. So much fun 😂 Great model release! https://t.co/afgSeYd0X7 → link

@sudoingX · 2026-08-08T09:19

people keep asking what an agent can do for them, book the flight, write the code, and that's the wrong question.

keep one around long enough and it stops being a tool. an agent that holds your training log, your ledger, your calendar, your 1am questions, becomes the most complete record of a human that has ever existed.

your diary only knows what you confessed. this thing knows what you did, what you earned, what you gave, what you asked when nobody was awake, and it cross references all of it and it never forgets.

now ask where that record lives.

if your people run on someone else's cloud, the deepest file ever assembled on you sits in a database you don't control, under a retention policy you didn't write, one breach, one subpoena, one acquisition away from being somebody else's asset. a staff that reports to their shareholder before it reports to you.

that's the real case for local ai, and it's why the box beside my monitor matters more than its tok/s.

last night an agent built and tested software for me and the whole loop ran on machines i own, model to keystroke, nothing left the room. that's the direction, your people on your metal, your life in files you can actually hold.

agents are going to become companions, that part's already decided. the only open question is whether you'll own yours or rent them. → link

Developer Tools & Agents

@Teknium · 2026-08-08T16:51

Hermes Agent Desktop keeps evolving → link

@Teknium · 2026-08-08T10:40

RT @iamlukethedev: Agent plugins should not be trapped inside one platform

Hermes now supports portable Agent Plugins v1 packages.

Instal… → link

@jxnlco · 2026-08-08T17:48

today, codex

  1. save me $2000 a year in canceled subscriptions
  2. found 44gb of extra disk space → link

@jxnlco · 2026-08-07T20:37

RT @sharifshameem: Siri sucks. So I made a way for Codex to act as my iPhone's voice assistant.

Now Codex can read my screen, control apps… → link

@kunchenguid · 2026-08-08T04:44

sigh.. i have to say something here

imagine you hired a new developer to your company, and on day one he did some terrible work, over-engineering your codebase, speaking jargons all day without context, making all kinds of “genuine mistakes”, lost trust with everyone around him

and then he goes on and tell you - in order for me to do a better job, you need to delete your entire company’s culture and workflow, and have everyone bend over to do things his way, only then can he do reasonable work

oh - and no one tells him how to do his job, NO ONE. he’s always the smartest one in the room and despite doing a terrible job on day one, despite the only results on his resume were vibe coded demos of games that already existed, he demands that you give him a big charter and let him go dark with no communication, taking no feedback

would you have hired a teammate like this?

when humans feel frustrated after using a model, let’s figure out how to RL the model better so they become a better teammate

don’t let a bad model RL you → link

@kunchenguid · 2026-08-07T21:47

alright i just did quite a few sessions with muse spark 1.2 and can share some initial insights

  1. i only used their contributor tier which has data collection but super cheap ($0.1/Mtok input, $0.002/Mtok cached input, $0.2/Mtok outupt)

even with this subsidized pricing, less than an hour of usage resulted in slightly over $1 of cost (fig. 1). if i do this all month it'll be more expensive than the highest Anthropic/OpenAI subscription

so unfortunately this means as a consumer it simply doesn't make sense to use it heavily. Anthropic/OpenAI's subsidized subscriptions are just too much value, making it practically impossible to choose anything else

  1. for enterprises who can't use the subsidized subscriptions from Ant/OAI, there's no way they would agree to data collection so they would pay the full price for muse spark

at full price, the ROI simply isn't there. see fig. 2 - on the deepswe benchmark, muse spark would cost a lot more than gpt-5.6-luna for worse results

  1. the model produces tokens very fast. with streaming out, my terminal scrolls so fast that i can't read anything until it stops

however, it seems to need more tool calls and steps to accomplish tasks which is cancelling out the raw output speed

  1. qualitatively, when comparing with grok 4.5 and opus 4.8, muse spark 1.2's instruction following capability is visibly weaker, especially at long context

when using it as my firstmate, it would admittedly skip instructions in system prompt and do things in its own way

however one thing i do like about the model is its response - it will format text responses nicely with sections and bullet points, making it quite easy to parse

so - overall i think Meta still has ways to go in terms of developing a competitive mainstream model

my suggestions to Meta:

  • given the model capability itself still seem behind, especially in a qualitative way, acquiring more human data seems crucial

  • more data depends on wider adoption, which would likely come from consumers not enterprises. however the contributor API pricing tier is not competitive enough in practice, so they really should consider offering a subsidized consumer subscription just like anthropic and openai. that's pretty much a ticket into the game at this point → link

@badlogicgames · 2026-08-07T19:27

i did it. i made sol go in thinking cirlcles forever. and all it took was 1 derranged design doc. https://t.co/gkOKIsbg8e → link

@thdxr · 2026-08-07T23:07

opencode2 has tool search and codemode built in

this means it can handle infinite mcps, custom tools, apis, etc

and it's the simplest implementation we've seen https://t.co/GgRJiIhaYX → link

@kunchenguid · 2026-08-07T19:38

"it caught 89% of dangerous commands"

well... what about the other 11%? being better than humans (who usually won't look carefully) doesn't mean it's useful

this proves reviewing each command is simply not the right solution

instead, we need to secure the environment and set boundaries. linking my previous tweet below where i shared my solution - → link

Infrastructure & Performance

@jezell · 2026-08-08T14:08

Wasmtime is a great WASM engine, but unlike V8 it doesn't have lazy / tiered JIT built in. Since it's primarily targeting server side loads it compiles the whole WASM up front when you load. If you can cache that, it's a one time hit, but lazy tiered JIT is a lot nicer for UI where you want fast starts.

So, I'm adding lazy tiered JIT to wasmtime. Initial testing shows about 219x faster startup / initialization on 10mb dart WASM files since it doesn't have to compile things until they are actually used. → link

@jezell · 2026-08-08T01:25

LibreOffice rendering on Flocker WASM with Skia Graphite took codex about 2hr. Dart port was 1.5 months with /goal in at around 50% of the codebase ported. Multi language is the future. https://t.co/WahscVO33j → link

@hnasr · 2026-08-07T20:30

Select For Update does a write in PostgreSQL.

In a Postgres heap page, each tuple has a header which provides metadata about that tuple.

When a transaction executes a select for update the tuples discovered as part of the predict will have their header updated to be marked as locked and an attribute xmax marks the transaction that did the lock.

Specially the two bits in the infomask header HEAP_XMAX_EXCL_LOCK and HEAP_XMAP_LOCK_ONLY are set.

This dirties the page and also is logged in the WAL. Which later may git shipped to a standby replica.

That dirty page later may be flushed to disk by the background worker.

On commit the bits are cleared.

Fascinating that a read operation can cause a bunch of write IOs. → link

@ivanfioravanti · 2026-08-08T13:16

Apple Container v1.2.2 is out and now there is the possibility to create a standalone K8S cluster that you can manage through kubeadm! 👀

https://t.co/lbILsRoyaV https://t.co/zT6MOxWvdD → link

@kunchenguid · 2026-08-08T17:57

PSA - quota-axi now has a human-friendly "--tui" mode that renders a real time dashboard of your quotas

its primary value is still to expose the data to your agents so they can make quota-aware routing decisions, but we humans need something nice too

npm i -g quota-axi https://t.co/Hke4XBViJm → link

Industry & Strategy

@swyx · 2026-08-08T18:28

reading thru applications. over 600 people applied, 100 admitted last night.

we are going to kill SO MUCH SAAS https://t.co/nrZO3rXkkU → link

@swyx · 2026-08-08T07:45

$10,000 kill my saas in a weekend competition is live!

TECH STACK: any coding agent any model up to $500 in token spend incl subscriptions

see luma for description. join waitlist if late to this - finish line extended to Wednesday.

the brief is live and people are prompting their clankers already. get to it!!! → link

@TheAhmadOsman · 2026-08-07T21:40

Top 5 moments for Opensource AI

  • Llama 3

  • Qwen 2.5

  • DeepSeek R1

  • GLM 4.5

  • Kimi K3

These are the moments that changed the trajectory of AI for all of humanity → link

@ivanfioravanti · 2026-08-08T08:29

What will happen when Chinese labs will release a model stronger than Western labs?

Distillation accusations won't be valid anymore.

Can't wait to see what the future holds for us. → link

@TheAhmadOsman · 2026-08-07T23:41

All panels and presentations from our Local AI Summit at AIE's World's Fair 2026 is now available to watch online

Watch us make Local AI The Default https://t.co/mLk6e6QQgj → link