Executive Summary
The dominant story of the last 24 hours is Ternary Bonsai 2, a compressed 27B reasoning model from PrismML that runs on consumer GPUs as small as 12GB VRAM while retaining ~98% of the original Qwen 3.8 27B's performance. Multiple developers independently confirmed it running full agentic loops with 262K context on RTX 3090s and even 3060s. Simultaneously, TypesafeAI's "Jev" model generated massive buzz as an ultra-fast inference model, though questions emerged about its moat and whether it requires a frontier model behind the scenes. Open-source and local AI advocacy intensified, with developers showcasing GLM-5.3 Flash, DeepSeek V4.1 Flash, and Qwen 3.8 Next Flash as viable alternatives to OpenAI/Anthropic offerings.
Key Events
- Ternary Bonsai 2 27B released — PrismML compressed Qwen 3.8 27B to ternary weights (1.75 bits, 5.95GB on disk), enabling it to run with full 262K context on an RTX 3060 12GB. Multiple users confirmed 50-70 tok/s decode on 3090/4070 cards with 128K+ context. → link
- TypesafeAI "Jev" model launches — Ultra-fast inference model generating significant mindshare, but critics note it may require a frontier model behind the scenes and faces immediate open-source reproductions. → link
- OpenCode adds
/btwfeature and intent-aware permissions — New quick-question feature avoids interrupting sessions; plugin system allows blocking agent shell commands without predicting every possible command. → link - Hugging Face Hub v1.32 deduplicates local cache — Identical Xet-backed files across repos are stored only once locally, reducing disk usage from ~100GB to ~20GB for shared weights. → link
- GLM-5.3 Flash self-improvement loop in production — Infra agent reads hand-tuned kernels, writes optimization skeletons, and feeds improvements back into the pipeline. → link
- ChatGPT desktop app adds Chrome extension support and multi-account plugin connections — Users can now bring browser extensions and multiple account contexts into ChatGPT. → link
- Pocket FM launches "Sherpa" AI writing assistant — Trained on 100M+ hours of listening data; helps writers develop characters, map seasons, and optimize dialogue for the audio fiction platform that grew from 21M to 500M ARR. → link
Analysis
Local AI is reaching an inflection point. Ternary Bonsai 2 represents a meaningful leap in model compression — a 27B dense model running with full context on a 5-year-old 12GB card changes the accessibility calculus. The simultaneous buzz around Jev (closed, fast) and Bonsai 2 (open, compressed) frames a central tension: proprietary speed vs. open accessibility.
Agentic tooling is maturing rapidly. OpenCode's /btw, intent-aware permissions, and codemode visualizations show agent infrastructure becoming production-grade. Amp's philosophy of "no tricks, just primitives" suggests the industry is converging on letting models use simple tools well rather than building elaborate scaffolding.
What to watch: Whether Bonsai 2's 98% retention claim holds under sustained agentic workloads (benchmarks pending); whether Jev's moat survives open reproductions; and whether the RTX 3090 secondary market (now $1,400+) continues climbing as local AI demand intensifies.
Tweet Feed
Local AI & Model Compression (Ternary Bonsai 2)
@victormustar · 2026-09-18T09:36
RT @xenovacom: NEW: Ternary Bonsai 2 just landed on Hugging Face — a 27B reasoning model in ternary weights, built for the agentic era 🤯 9… → tweet link
@sudoingX · 2026-09-18T16:56
two months ago prismml crushed qwen 3.6 27b to 3.9gb and i ran it on a 6gb 1660 super, an 8gb 3060 ti and a 3090, and the 8gb card held a full agent loop. bonsai 2 is the same trick on qwen 3.8 27b, 5.9gb, and their own table puts it a hair above the full precision qwen 27b i spent all summer calling the king of the 3090. their demos ran on a 5090. i am pulling it onto an rtx 3060 12gb right now, the most owned gpu on steam, then a 16gb card with the vision tower on, then the 8gb question. receipts tonight, filled context and a real agent loop, not a fresh context tok/s screenshot.🧵 → tweet link
@sudoingX · 2026-09-18T17:22
loaded. bonsai 2 27b is sitting on an rtx 3060 12gb right now, served through the prismml llama cpp fork with the full 262k context window resident, 11.7 of 12 gigs in use, hermes agent pointed at it. first receipt before a single benchmark: a 27b holds its entire native window on a 12gb card. now we find out what it does and how fast it runs. for anyone new here, bonsai 2 is not a new model. prismml took qwen 3.8 27b, a dense 27b that needs around 16gb at q4, and compressed the weights to ternary, three values per weight, 1.75 bits each, 5.95gb on disk, and their table says it keeps 98.2% of the full model on their suite. the 3060 is the number one desktop gpu on steam at 3.92% of every pc, a five year old card, so if a 27b runs as an agent here, it runs for more people than on any other card on earth. for anyone who followed the july run, bonsai 1 was the same idea on qwen 3.6 27b at 1 bit, 3.9gb, and it held a 30 minute agent loop on an 8gb 3060 ti. bonsai 2 is a bigger file on a better base, so the question moved from does it run to how much of the full model survived. tonight, in order: speed at empty context and at 64k filled, then a real build through hermes agent, the same spec the july run had to pass. receipts as they land, tool named on every number. → tweet link
@gospaceport · 2026-09-18T04:54
RT @DeathscytheUltr: I just setup Bonsai 2 27b 2bit on my 4070 12gb GPU using 163,840 context with Q4_0 K/V at 50.27 tok/s TG and 704.33 to… → tweet link
@gospaceport · 2026-09-18T06:31
Here this chart is much clearer and TIL llama.cpp logs cant be read like vllm logs. Still, 61K in 60s all layers offloaded on a single 3090. It's fast. Overall quality, I'm still out on. I am head to heading it on the same prompt Qwen 3.8 FP16. Results will be head 2 head playable for you whenever it finishes... I did notice it not handling subagents well and them getting stuck in read loops. Could be unrelated, but occurred twice. Running as just a main thread agent now and it is at least sailing along. → tweet link
@gospaceport · 2026-09-18T02:45
Cat Walking on a Fence Animated SVG - Ternary Bonsai 2 27b - 24K toks, 64.4 tps decode, 1x 3090, 128K ctx → tweet link
@TheAhmadOsman · 2026-09-18T00:11
9x size reduction from Qwen 3.8 27B while maintaining 98% of the performance. The future of Local AI is so good you won't believe it → tweet link
@TheAhmadOsman · 2026-09-18T00:42
234 days later, we got Opus 4.5 / 4.6 intelligence and capabilities running on a single RTX 3060 with 8GB VRAM → tweet link
@gospaceport · 2026-09-17T22:42
A single 3090 can now enjoy DENSE model smell 🤓 → tweet link
@gospaceport · 2026-09-17T21:57
Bad news on "cheapest used" 3090 front from eBay, now $1400 new lowest priced. $2400 for 4090 and 5090.... just forget about those RN. → tweet link
TypesafeAI "Jev" Model
@jezell · 2026-09-18T04:36
Does Jev have no moat? They have the name brand at least. Everyone is making a Jev knockoff now. Better acquihire while the name is hot. → tweet link
@jezell · 2026-09-18T01:34
So who is gonna acquire @typesafeai? They seem to have gotten a lot of mindshare pretty quickly, but sounds like they still require a frontier model behind the scenes to work. → tweet link
@levelsio · 2026-09-17T21:04
Why is everyone talking about Jev today → tweet link
@iamdevloper · 2026-09-17T21:17
Ok Jev, let's take you for a whirl → tweet link
@thdxr · 2026-09-17T23:09
should we put jev in our free tier → tweet link
@RayFernando1337 · 2026-09-17T21:36
RT @alexlavaee: Jev explained in under 10 min: what it is, how parallel constrained decoding works, when to use it, and when not to. Shout… → tweet link
@badlogicgames · 2026-09-17T19:37
RT @Bewinxed: I made @typesafeai 's new ultra fast model, Jev, generate text, even though it shouldn't, that's fine because I can't read, a… → tweet link
@swyx · 2026-09-18T01:01
RT @aiDotEngineer: congrats to Diogo on a full launch! for more on typesafe: check his aie talk: https://t.co/pPg9t1cIk8 out now! https:/… → tweet link
Open Source / Local AI Advocacy
@TheAhmadOsman · 2026-09-18T04:15
My current LLMs stable - GLM 5.3 Flash - DeepSeek V4.1 Flash - Qwen 3.8 Next Flash - Qwen 3.8 27B. Just a year ago you wouldn't have believed this kind of performance was possible outside of OpenAI & Anthropic. But we're just getting started and this is the worst it'll ever be → tweet link
@TheAhmadOsman · 2026-09-18T03:47
Opensource AI must win. Local AI will be the default. Freedom of intelligence is a must. → tweet link
@TheAhmadOsman · 2026-09-18T05:30
"We're gonna make Local AI as accessible to the average Joe as the price of these bananas" — @OsmanticAI, making AI banana-priced → tweet link
@TheAhmadOsman · 2026-09-18T04:58
Playing with a few RTX PRO 6000s and a DGX Station tonight. I feel privileged and that is actually a problem. In a few years, I genuinely believe we won't need such massive amounts of compute to run quality models --- this is the problem we are building @OsmanticAI to fix → tweet link
@RayFernando1337 · 2026-09-18T14:01
Subscriptions are dead...local models are smart enough to do things on device and most people are overpaying to do basic things. I'm going to break the trend and I'm launching SayRay to solve a problem for my friends who want to see text when they speak. → tweet link
@alexinexxx · 2026-09-17T20:25
don't get me wrong i love local AI but spending thousands on hardware to become your own worse cloud provider is an incredible hobby → tweet link
Agentic Coding Tools (OpenCode, Amp, Claude Code)
@thdxr · 2026-09-18T00:42
in the next version of OpenCode we added /btw which is a long awaited feature. it lets you ask a quick question without interrupting the current session - i use it to get status updates on long running work → tweet link
@thdxr · 2026-09-17T23:08
RT @OpeOginni: Built an intent-aware permissions plugin for @opencode. No need to predict every shell command an agent might use. I blocked… → tweet link
@thdxr · 2026-09-18T01:00
more codemode fun, doing a large mysql operation and the agent wanted to do it 50K rows at a time. it writes a single codemode script which runs it in a loop and then you visualize it step by step even though it's code → tweet link
@sqs · 2026-09-18T07:56
This is exactly what we want to achieve with Amp: get some simple primitives right (like threads and orbs and runners), and let the models use them. No tricks. No "must-have" skill stacks. And it means Amp gets better when the models get better. → tweet link
@sqs · 2026-09-18T08:08
RT @jkudish: It is time... I am simply no longer interested in AI subscriptions that don't work inside @AmpCode → tweet link
@jxnlco · 2026-09-18T18:11
RT @trq212: We're adding support for AGENTS.md to Claude Code. Starting today in version 2.1.277, if there is no CLAUDE.md in a folder, C… → tweet link
@steipete · 2026-09-18T16:48
gog has an mcp server now! → tweet link
@steipete · 2026-09-18T16:46
Collaborating on a session is so good. Even see what someone's typing to prevent sending the same to the agent. → tweet link
@thdxr · 2026-09-18T04:22
i had fable doing some inference optimization work, it was making a ton of progress then it realized what it was doing and started to reject my requests so i switched to astra and told it to spawn fable subagents to do the work back in business → tweet link
@thdxr · 2026-09-18T05:08
a few weeks ago i had OpenCode train a small local model to classify shell commands into a few categories to solve a UI issue. obviously now jev can do that but it'd still work better local so... → tweet link
Model Releases & Performance
@ivanfioravanti · 2026-09-18T08:06
DwarfStar DeepSeek V4.1 Flash: this time I will work on the existing PR from super @kernelpool. Combined effort FTW 🙌 → tweet link
@ivanfioravanti · 2026-09-18T03:48
MiniMax H3 is becoming faster week after week 💪 → tweet link
@louszbd · 2026-09-17T19:15
We were bringing up the inference stack for GLM-5.3-Flash. And sitting there watching agent work, we found that what it got back after a change mattered about as much as how good the model was. That's what we mean by dense feedback. Our infra agent runs on GLM-5.3. It read through kernels tuned by hand and wrote down what it found as optimization skeletons. Whatever held up went back in so the next kernel takes less work. Early days, but it's already running in production. Different labs put self improvement in different place. Here's a small piece of ours from our everyday work. → tweet link
Hugging Face & Infrastructure
@victormustar · 2026-09-18T15:50
Hugging Face cache is way more than a folder full of downloaded models 🤫 With huggingface_hub v1.32, identical Xet-backed files across repos are stored only once, so if 5 repos share the same 20 GB weights file: -> before: ~100 GB on disk -> after: ~20 GB. And once the blob is local, another repo can reuse it with zero payload download 💃 Xet already made the remote storage content-addressed. Now that identity extends into your local cache too. → tweet link
@victormustar · 2026-09-18T14:15
RT @wildmindai: HOT! LTX-2.5 Video/Image Enhancer LoRA. Updated to V2! generative quality restore for old/low-res images-videos - runs fas… → tweet link
AI Tools & Applications
@LinusEkenstam · 2026-09-18T18:55
This app is wild. They went from 21M to 500M in ARR in 3 years. Now they are introducing Sherpa. It's trained on 100M+ hours of listening data, including where audiences drop off and spend coins to unlock another episode. Whats insane is that US listeners are spending +2 hours a day in the app, while spending less than 1h a day on TikTok. Sherpa is a new hot take on AI assisted writing. You bring an idea. Sherpa helps develop the characters, map the season, work on dialogue, and figure out where to leave people hanging. I like the idea of a writer being able to work with that feedback while building their own story. Still need taste. Still need a reason for anyone to care. And I'd want this to leave room for the weird ideas that don't look like yesterday's hits. Pocket FM is helping more people actually finish the story they've had in their head for years is a pretty good use of AI. Really want to see what people make with this. → tweet link
@MengTo · 2026-09-18T12:38
I made a learning site for grades 1–7 with books for math, English, and science. My kids loved it. I used GPT-6 Astra with the Hyper3D Rodin MCP to create the 3D shelf souvenirs, candles, lamps, and pins. What stood out was the speed and quality. Rodin gave me richly detailed, high-quality, production-ready assets for the site. You can see the detail in the metalwork on the candelabra and lantern. → tweet link
@jxnlco · 2026-09-18T17:17
RT @mxstbr: Starting today, you can connect multiple accounts with most plugins in ChatGPT! 🎉 Bring context from your work and personal a… → tweet link
@jxnlco · 2026-09-18T17:30
RT @JamesZmSun: Today, we're launching support for Chrome extensions in the ChatGPT desktop app! Bring the extensions you use every day to… → tweet link
@jxnlco · 2026-09-18T00:04
RT @ChatGPT: One of the most underrated features in the ChatGPT desktop app: Appshots. Appshots take the context on your screen and bring… → tweet link
@gdb · 2026-09-17T21:31
Astra for Law — tools and skills to radically accelerate law firms. Includes data privacy as a core feature, and 26 partner-built plugins and 47 community plugins for legal work in ChatGPT. → tweet link
Benchmark & Research
@jezell · 2026-09-18T07:58
RT @EpochAIResearch: Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verif… → tweet link
@jezell · 2026-09-18T03:35
RT @S1r1u5_: On July 25, we hacked OpenAI. Two bugs let us take over ChatGPT/Codex accounts of OpenAI employees (+some unaffiliated users)… → tweet link
Developer Commentary & Open Source
@iamdevloper · 2026-09-18T16:36
there is still no agreed upon icon for an Agent or an LLM, like there is for phone, mail or a file → tweet link
@iamdevloper · 2026-09-18T13:50
Give a man a fish and he eats for a day. Give a man an open source library and he maintains it unpaid for a decade while a trillion pound company quietly builds its entire business on top of it. → tweet link
@badlogicgames · 2026-09-18T00:42
hi kids, gramps here. learn your algorithms and datastructures, cause SOTA frontier models will still fuck you up bad behind your back, especially if you don't read your code anymore. → tweet link
@thdxr · 2026-09-17T22:15
the last 20% to get a product shipped and successful is as difficult as ever. this is where 99% of programmers have always fallen off. the first 80% is a lot easier and so we're seeing more abandoned projects than ever → tweet link
@ivanfioravanti · 2026-09-18T02:40
All these posts by AI labs about slowing down to improve security is nonsense and must stop. Do it and stop writing about this. You (labs) are responsible for the product released, if they are not secure enough, you should slow down and don't release. If you make mistakes and cause damage during your tests, you are responsible, no one else. → tweet link
@jack · 2026-09-17T22:57
RT @saranormous: "Why would I look at code? It's like assembly, like a compiled artifact." Sensei @karpathy fully agentic, 18m after "I do… → tweet link
Apple & Hardware
@ivanfioravanti · 2026-09-18T08:20
Meeting at Apple Developer Center in Shanghai done! 🚀 Ready to push Apple Silicon Local AI to the next level together! → tweet link
@FrameworkPuter · 2026-09-18T16:29
Ever wanted to have an 11 hour video conference with your colleagues and friends? Now you can! In our latest battery life test video, we're showing two Framework Laptop 13 Pro's reaching 11 hours of battery life on a Google Meet call. → tweet link
@MatejKnopp · 2026-09-17T21:08
Flagship Windows laptop, installing OpenSSH server. This has been going on for like 15 minutes now. It's not stuck, it's progressing. Slowly. This used to take seconds. I guess tokenmaxxing at microsoft is going great. → tweet link