← Tech / AI / IT Monitor Index Tech / AI Generated 2026-08-26 19:31 UTC

Tech / AI / IT Monitor

August 26, 2026 · Based on tweets from the last 24 hours · 219 tweets analyzed · model: ollama-cloud/glm-5.2:cloud

Executive Summary

The last 24 hours saw major releases in open-weight models, notably GLM-5.3-Flash and Qwen3.8-Flash-Next, driving a massive surge in Local AI capabilities. Apple's new M5 Ultra Mac Studio was widely celebrated for enabling frontier-class models to run locally with unmetered tokens, challenging incumbent AI hardware like Nvidia's DGX Spark and Cerebras' newly announced CS-5. Developer tools are rapidly evolving to support this local-first shift, with integrations like Ollama's seamless Claude Desktop support and new agent harnesses like Hermes and Amp gaining traction.

Key Events

Analysis

The overarching trend is the aggressive shift towards "Local AI." Open-weight models (GLM, Qwen) are reaching intelligence parity while optimizing active parameters, perfectly timed with Apple's massive hardware upgrade (M5 Ultra) and Framework's desktop push. Simultaneously, inference startups renting cloud GPUs are facing existential threats as local, unmetered compute becomes viable at API speeds. On the software side, developer tools (Claude Code, Amp, Hermes, Ollama) are heavily integrating Model Context Protocol (MCP) and sophisticated agent harnesses to automate complex workflows. Watch for Nvidia's response to Apple's silicon and OpenAI's custom chips, as well as whether cloud inference providers can lower margins enough to retain developers.

Tweet Feed

AI Models & Open Weights

@Prince_Canuma · 2026-08-26T18:33

Congrats to @Zai_org on GLM-5.3-Flash! Day 0 support in @Nativ_AI 🎉

320B params, 18B active · 1M native context · fully on-device.

On an M3 Ultra (512GB) with Nativ v0.3.5 and 4-bit you should get:

⚡ Up to 505 tok/s prefill · 32 tok/s decode 🚀 Batch 4 lifts decode to 70.9 tok/s 💾 Peak under 380GB, fully resident, no offload 📏 Benchmarked to 128K context

A 320B MoE running entirely on your Mac.

Try Nativ 👇

https://t.co/nu6v5JIVBe → tweet link

@louszbd · 2026-08-26T17:24

RT @ZixuanLi_: We’ve updated the chat template for GLM-5.3-Flash. If you downloaded the model before this update, please download it again.… → tweet link

@alexinexxx · 2026-08-26T16:33

RT @Zai_org: Introducing GLM-5.3-Flash

  • Leading capabilities at a highly competitive price
  • Natively multimodal with a 1M-token context… → tweet link

@TheAhmadOsman · 2026-08-26T16:21

GLM 5.3 Flash and Qwen 3.8 Flash Next are great examples of Local AI progression

This is the good timeline → tweet link

@ivanfioravanti · 2026-08-26T16:18

I know it's not a fair comparison, but I collected benchmarks published for GLM 5.3 Flash and Qwen 3.8 Flash Next.

Second one much faster, first one more intelligent. Combo/Hybrid usage will win as always. https://t.co/TPazavBDfS → tweet link

@ivanfioravanti · 2026-08-26T16:16

RT @UnslothAI: Qwen3.8-Flash can now be run locally! 🔥

The 125B MoE model outperforms Claude-Opus-4.6 (Max).

Run on 75GB RAM via Unsloth… → tweet link

@Prince_Canuma · 2026-08-26T08:11

.@Alibaba_Qwen announces Qwen3.8-Flash-Next (Qwen 4 architecture), a new open-weight multimodal MoE model.

The model will be released today and will come with Day-0 support on @Nativ_AI → tweet link

@victormustar · 2026-08-26T14:01

here we go again: deployed a FREE public endpoint for Qwen3.8-Flash-Next 🚀 (going at +100 tok/s)

No token needed, OpenAI-compatible, vision + tool calls, 262K context, thinking from xhigh → off. Light rate limiting, be nice to your neighbors 🤗

4× H200 · FP8 · SGLang cookbook · ~140 tok/s per stream · ~100 tok/s @ 16 concurrent · 0.8s TTFT

Guide + chat UI 👇 → tweet link

@ollama · 2026-08-26T14:28

glm-5.3-flash will be available shortly on Ollama's cloud. → tweet link

@Teknium · 2026-08-26T14:53

Ox Alpha free period is over, but GLM-5.3 Flash is now available if you'd like to continue using it! https://t.co/7tqjDeLX4b → tweet link

@LinusEkenstam · 2026-08-26T10:43

Ox Alpha is GLM(-5.3 Flash) by zAI.

I hate to say I told you so… → tweet link

@KingBootoshi · 2026-08-26T05:43

reading kimi-k3's chain of thought is incredibly insightful → tweet link

@TrungTPhan · 2026-08-26T03:08

RT @petergyang: Today, I’m open-sourcing /fuck-cancer, an AI skill that helps patients and caregivers navigate cancer diagnosis and treatme… → tweet link

AI Hardware & Chips

@alexocheema · 2026-08-26T18:56

RT @twid: Callout for @exolabs spotted 👀 https://t.co/wzUirT5gEC → tweet link

@MilksandMatcha · 2026-08-26T18:39

But has anyone really THOUGHT DEEPLY about why after 8 months NVIDIA and Groq are going out with an old, really small, model... → tweet link

@TheAhmadOsman · 2026-08-26T18:08

To be clear, CUDA is still more mature than MLX if you’re comparing GPUs to the new M5 Ultra Mac Studio

However, the ecosystem took years for CUDA to mature and we need competition and people to optimize all available hardware

This is a win for Local AI in anyway you look at it → tweet link

@MilksandMatcha · 2026-08-26T16:07

Last week, Cerebras CTO @seanliecs announced CS-4, 30x faster than the GPU.

This week at #hotchips2026, Cerebras announced CS-5, another step function faster than anything we've seen before with up to 10,000 tokens/sec/user on models such as Gemma 4 31B and gpt-oss-120B.

Seana and I will be doing an AMA on Cerebras, CS4/5/6, Hot Chips announcements. Let us know what questions you have. → tweet link

@ivanfioravanti · 2026-08-26T15:22

Previously previewed as Ox Alpha! 100T tokens per day. Running entirely on Chinese AI chips.

Let that sink in https://t.co/TluSPMLGgb → tweet link

@jezell · 2026-08-26T14:06

RT @SemiAnalysis_: OpenAI Jalapeño: Better Than Nvidia Blackwell OpenAI’s self-designed ASIC compared with Rubin, Jalapeño’s TCO, throughpu… → tweet link

@alexinexxx · 2026-08-25T23:40

RT @PatrickToulme: OpenAI Jalapeño is truly the first AI silicon developed by GPT-Astra and other OpenAI internal models.

I believe they… → tweet link

@alexocheema · 2026-08-26T14:38

Apple is upgrading customers to the new M6 Mac Mini for free.

Email from Apple:

"Thank you for your recent Mac order. We know you're looking forward to receiving your purchase. As you may know, Apple recently announced the new Mac mini. Since your order has yet to ship, we automatically upgraded you to the new Mac mini at no additional cost." → tweet link

@alexocheema · 2026-08-26T10:59

10-12 week lead time now on 256GB M5 Ultra Mac Studios.

The demand is insane. Secure your compute. https://t.co/7Zbe2Ofjer → tweet link

@TheAhmadOsman · 2026-08-25T19:40

The DGX Sparks got killed today by Apple → tweet link

@sama · 2026-08-25T19:53

we made a chip and it is fast → tweet link

@tinygrad · 2026-08-26T02:43

Can we get a price reveal for the MI350P? If this is priced competitively with RTX PRO 6000 Blackwell and gets a normal 3 fan GPU cooler, it could be a real winner for @AMD. @AnushElangovan https://t.co/QaVme0PMaC → tweet link

@gdb · 2026-08-26T05:41

ai for chip design is underrated → tweet link

Developer Tools & Infrastructure

@ollama · 2026-08-26T03:26

Ollama v0.33 is here!

You can now easily configure Claude Desktop to seamlessly work with Ollama as a third-party gateway provider.

One toggle. Cloud & local models just work👇 https://t.co/NUPYfWdjW3 → tweet link

@thdxr · 2026-08-26T17:10

in latest version of opencode2 we include TPS indicator

we can finally calculate it accurately! it's surprisingly tricky https://t.co/MdmONUA1eA → tweet link

@thdxr · 2026-08-25T21:11

in the next version of opencode2

you can connect multiple accounts for the same provider and switch between them

if you use this against me i will be mad https://t.co/1m3mieHP78 → tweet link

@sqs · 2026-08-25T19:38

Amp terminal (uses ghostty-web) in orbs and runners has a mobile accessory keyboard now

I laid out the keys myself...was fun. Anything missing? ^ and Meta let you tap another key of course. Could have configurable shortcuts. https://t.co/QzlYbGKFVg → tweet link

@sqs · 2026-08-26T02:31

For everyone who: • doesn't want to commit .agents/setup OR • needs to run tailscale up before repo cloning

Amp's now for you, too. → tweet link

@jezell · 2026-08-26T17:31

@tiptap_editor is solid. hocuspocus + yjs is a really nice combo for realtime collab. https://t.co/ZQHAZd0iFw → tweet link

@jezell · 2026-08-26T17:23

RT @kubernetesio: Kubernetes v1.37: Garhwal - https://t.co/nKo1DDc6Gr #Kubernetes → tweet link

@jezell · 2026-08-26T17:13

RT @theinformation: ClickHouse has surpassed $350 million in annual recurring revenue, up 40% since May, as demand from AI agents accelerat… → tweet link

@jezell · 2026-08-26T17:30

RT @jonathangrahl: Github using Vitess 👀 https://t.co/XAVbNMKyMc → tweet link

@jezell · 2026-08-26T13:39

RT @mitchellh: The Ghostty memory usage memes are dead with the upcoming 1.4 release, much to the dismay of internet trolls everywhere. ~10… → tweet link

@swyx · 2026-08-26T06:02

PSA: do not use codex "locked use" capabilities right now. it is currently relying on unstable mac features and has completely locked me out of my macos keychain twice this week.

thx @_chenglou for linking to apple developer forums acknowledging this is a "known bug". just avoid. ofc, would be nice to do everything in cloud, but cloud isn't there yet. → tweet link

@thdxr · 2026-08-26T00:57

i complained a lot about the mcp spec in the past but it's getting very good now → tweet link

@kunchenguid · 2026-08-25T20:33

Anthropic is again refusing to follow industry standards like AGENTS.md, and insists that CLAUDE.md needs to be special

now here's what's wrong with their stance:

  1. the underlying attitude from Anthropic is that they know better than everyone else

OpenAI is doing it wrong. the open source community is doing it wrong. everyone is doing it wrong

only we anthropic know the best way, and we will not budge → tweet link

@steipete · 2026-08-26T06:38

RT @pauldix: I think the release of Bun 1.4 marks the end of programming as we know it. Anthropic and OpenAI developers live in the near fu… → tweet link

@Teknium · 2026-08-26T04:10

RT @HermesWatcher: This is a ridiculous integration drop for Hermes.

44 new remote MCPs just landed in main, including Canva, Dropbox, Git… → tweet link

@thdxr · 2026-08-26T13:30

our anti fraud system gets access to honeycomb, stripe, and database to correlates things and find very obvious fraud

the models are incredible at doing this via codemode. the whole script didn't even fit in the screenshot

composing these tools together to get what it needs https://t.co/mtvEh2wdS1 → tweet link

@swyx · 2026-08-26T01:08

RT @StephanEwen: Something we @restatedev are incredibly proud of is powering Replit's agents with durable execution.

Their scale and the… → tweet link

Local AI & Hardware Clustering

@alexocheema · 2026-08-26T04:20

Apple's marketing video posted by Joz, includes the 4 x M5 Ultra clustered solution. The cluster has 2TB unified memory @ 4.8TB/s.

"Run trillion parameter frontier models locally" https://t.co/qqmXMr3H4X → tweet link

@alexocheema · 2026-08-25T21:40

2 years ago, we achieved the first big milestone with @exolabs.

We clustered 2 MacBooks to run Llama 405B.

It felt like magic.

The consensus was running this model was only possible in a data center.

We ran it on consumer hardware, on 2 M3 Max MacBook Pros.

Most people thought it was a gimmick. It only ran at 2 tok/sec!

But, we believed that improvements to the software, hardware, and models would all compound.

So that maybe in a few years, we thought, this would improve 10x in software, 10x in hardware, 10x models = 1000x.

That was the vision.

We imagined a world where you would have frontier intelligence running quietly on your desk.

Today is the day that vision became reality.

The M5 Ultra is a 10x step-change improvement vs the M3 Max we originally clustered.

That, compounded with software improvements like RDMA over Thunderbolt, MTP and better kernels, and high intelligence density models like Qwen 3.8 27B, means we now have 1,000x better Local AI than when we started.

I am so grateful to the small group of people at Apple (including @doogie69 @awnihannun @angeloskath @DiganiJagrit @doogie69) who believed in this vision and had the foresight as well as the courage to take a swing at this early on. I'm confident they're just getting warmed up (looking forward to 4-bit / 8-bit compute units in M7 Ultra🤞).

With the M5 Ultra Mac Studio, we are going to have unmetered tokens running at API speeds at effectively zero marginal cost, running on your desk, so quietly and consuming so little power you won't even notice it.

Local AI is good now. → tweet link

@alexocheema · 2026-08-26T15:08

One of the great things about Apple embracing Local AI is that they have a track record of genuinely caring about privacy.

Jobs cared deeply about privacy. It's in the DNA of the company.

They have resisted pressure from nation states to build backdoors into the iPhone. As AI becomes more useful, it also becomes more intimate and invasive.

Local AI is our defense against mass surveillance and a handful of companies owning the intelligence layer, shifting control back to individuals.

Apple is also not exposed to the "AI bubble" e.g. data center buildouts or big training runs. Their AI capex has stayed relatively flat. If most inference moves local, they sell more devices, and win. → tweet link

@alexocheema · 2026-08-25T22:09

getting a lot of DM’s asking what hardware to buy.

we will put all the data on https://t.co/LtX5bSseMm as soon as we can.

we test every popular model that fits, to get an intelligence/speed pareto frontier for each device. then we calculate a score based on speed, intelligence and cost for each device. it’s the single best number that tells you what the best value hardware is based on what the experience of using it for Local AI will actually be like. → tweet link

@alexocheema · 2026-08-26T02:20

exo is currently top of r/LocalLLM (in response to this tweet).

answering questions there in detail on my alt: Longjumping_Crow_597.

feel free to ask any questions, exo or local AI related. link: https://t.co/ibCRA6AMNo https://t.co/aVONUdzgQh → tweet link

@FrameworkPuter · 2026-08-26T17:13

It's an excellent day for running open weights models locally! Qwen3.8-Flash-Next fits in its entirety at 4-bit on a 128GB Framework Desktop, and at 8-bit on the upcoming 192GB Framework Desktop. → tweet link

@sudoingX · 2026-08-26T08:40

i used my 2x dgx spark for a week and 99% of my workload is now local. all of it.

for a solo builder this is the setup. i stopped reaching for the cloud without even noticing. if you're on the fence about a second one, this is your sign. go.

and the mac mini crowd, it's a great machine, but if you live in the ai stack, cuda is the whole ballgame. choose wisely. → tweet link

@alexocheema · 2026-08-26T03:29

RT @atomic_chat_hq: Run Qwen 3.8 27B MLX locally on a 16GB Mac 💻

We released vision MLX quants on Hugging Face from 29.5GB down to 11.8GB… → tweet link

@jezell · 2026-08-26T15:34

@FrameworkPuter needs to start offering a big beasty server machine. Let's call it the Framework Build. A server grade machine that can function as a build runner for a low velocity team, or a personal build server for people who now are exceeding the limits of their dev boxes due to AI acceleration. Bonus points if it has docking stations for mobile devices. → tweet link

@gospaceport · 2026-08-25T20:57

Nvidia will respond if they are smart. This new Apple lineup is serious beef. I am not a mac guy, but facts are facts. → tweet link

@jezell · 2026-08-26T18:16

Like I've been saying, but the missing focus is solving this problem locally. Cloud is awesome for shared resources, but your sandbox provider isn't gonna give you iphones and macs. Not a problem if you only need a linux VM, but a lot of people need more than a linux VM. https://t.co/TwlZDkp7oH → tweet link

@alexocheema · 2026-08-25T23:35

If you are asking yourself how much value there is in ds4, my answer is a question and a matter of fact: 1. What problem does it solve that existing inference engines fail to solve? The fact it doesn't seem to solve any problem, and lacks any differentiating features may look like a limitation, but I'm not sure it is, and is also kinda of good design but also a statement of the limitations. This is not yet an inference engine that is solving a real problem. To use it and to fragment the ecosystem would be futile, the differentiator is insufficient. → tweet link

@NaderLikeLadder · 2026-08-25T22:45

Me: "lmao you can lease a Mac mini"

Alec: "that's just cloud with more steps" → tweet link

AI Agents & Workflows

@MilksandMatcha · 2026-08-26T16:55

Sherif asked Codex to buy the parts for a new computer and send them to his house. it did.

Then it noticed he accidentally ordered two motherboards, emailed the supplier, got him a refund, and found coupons.

@0xSero didn’t ask it to do any of that.

he also has Codex:

keep track of green card paperwork, taxes, and nonprofit filings edit videos locally with FFmpeg summarize his inbox check for approvals and continue with the next step turn repeated work into skills cheaper models can reuse run experiments overnight while he sleeps

In this episode of 'Independent Studies', Sherif walks through exactly how he setup his tiny staff of agents and the creative ways that he uses AI to make his life cheaper and more productive. → tweet link

@KingBootoshi · 2026-08-26T16:58

damn grabbing open sourced tools and instantly modding the fuck out of them w agents is crazy

just grabbing a tool thats already 80% there, agents get a VERY strong base to work off, then prompting for customizing the rest is extremely easy

i think proper etiquette here is making sure you create a fork of whatever your modifying, keep your changes on your repo and contribute PRs to any obvious headaches you come across (at this point contributing a PR off a fork is a quick prompt)

messing around with annotate OS repo, making it fit my workflow, it's actually incredibly fun to not start from scratch → tweet link

@KingBootoshi · 2026-08-26T09:46

single best tip i got for UX/UI design with agents, is always ask for multiple variations

by doing this you exhaust probability space with a variety of bangers

like one? anchor to it, generate 5 more probability paths

tune your style and ship it instantly. gg https://t.co/vj2BpNEMqJ → tweet link

@KingBootoshi · 2026-08-26T04:38

it’s good to be lossy with ai agents

meaning “be vague”

leave your answer up for interpretation

they’ll know what you mean more rather than if you give direct instructions for a task you don’t really know anything about

the point of a parable is to explain a foundational lesson that can be applied throughout all information flow

LLMs interpret lossy info better than details! → tweet link

@KingBootoshi · 2026-08-26T05:56

ADVISOR STEERING IS SO NICE AND NECESSARY

SAVES SUCH A BIG HEADACHE

CLOCK FABLE'S ASS LUNA !!

it's so good to pit LLMs against each other

because they anchor to different vectors they are less biased on keeping them in check

instead of getting tricked by itself

because the two different LLMs have completely different variables in their mathematical formula

it's truly a great way to anchor models to reality. they just have extreme ADHD! they need a body double to lock in ! → tweet link

@KingBootoshi · 2026-08-26T05:33

I just discovered the most BADASS open sourced tool for annotating your screen to communicate with vision agents MUCH more effectively

i cloned it and made my own modifications for personal taste but basically i hit cmd+shift+a and it activates a paint brush so you can draw on your screen, add rectangles (kinda like photoshop)

then you can screenshot the annotation and send it to your agent, so you can communicate information more precisely through annotated images! → tweet link

@jxnlco · 2026-08-26T04:21

RT @nickbaumann_: In case you haven't been paying attention, ChatGPT Work (web) now has the following primitives:

@gdb · 2026-08-26T01:55

triggered tasks in chatgpt: → tweet link

@gdb · 2026-08-25T22:19

you can now have chatgpt work securely sign into accounts: → tweet link

@gdb · 2026-08-25T20:47

by popular demand, $100 business seat now available in chatgpt! no 5-hour limit, more usage, etc.. → tweet link

@jsuarez · 2026-08-26T17:04

RT @jeffclune: Excited to share “AI Finds a Way.” 🦖 🦕✨ 🤖

AI can be surprisingly creative, outsmarting the researchers who use it. That can… → tweet link

@Teknium · 2026-08-26T06:07

So many useful things being discovered with Hermes Desktop's HUD Mode! → tweet link

@MilksandMatcha · 2026-08-26T17:42

We've reached the point where inference is so fast that tool calls are now the sore thumb in the equation

Those who think inference speed doesn’t matter because tool calls are slow are dumb

That’s like saying faster internet doesn’t matter because websites still have loading animations → tweet link