Executive Summary
Anthropic dominated the tech conversation with the launch of Fable 5.1 and Mythos 5.1, delivering major jumps in coding and automation benchmarks alongside a 75% reduction in cache read costs. Agentic engineering continues to mature as developers identify structural limitations—such as the inability for AI to judge "how it feels"—prompting the creation of robust verification, state-persistence, and multi-agent orchestration tools. In hardware, Apple leadership transitions are heating up with John Ternus taking a prominent public role, while memory innovations and aggressive model quantization (like Tencent's Sherry method) push local inference efficiency to new heights.
Key Events
- Anthropic launches Fable 5.1 and Mythos 5.1, featuring a 1M token context window and massive benchmark improvements in terminal automation, alongside a 75% price drop for cache reads. → link
- Ollama rolls out transparent per-token pricing for its Pro, Max, and Team plans, bundling monthly usage credits. → link
- Teknium releases Hermes Agent v0.21.0, introducing Bots Mode, Agent 2 Agent Comms, and persistent multi-gateway connections. → link
- Apple signals a potential leadership transition as John Ternus updates his X profile, prompting industry bets on Apple's future hardware and AI direction. → link
- Tencent's "Sherry" quantization method compresses a 1.5TB model to 214GB using 1.25-bit weights with minimal accuracy loss. → link
Analysis
The industry's focus is shifting from raw model scale to operational efficiency and orchestration. Anthropic's drastic reduction in cache read pricing signals a push to make long-running agents economically viable. Simultaneously, developers are vocalizing a structural "wall" in agentic engineering: AI struggles with subjective human experiences ("how does it feel"). To counter this, engineers are building highly resilient local fleets (e.g., FirstMate) and strict verification/advisor agents to maintain code quality. On the hardware side, vertical memory stacking and advanced quantization are making heavy local AI inference increasingly accessible on consumer hardware.
Tweet Feed
AI Models & Releases
@MilksandMatcha · 2026-09-01T18:43
wake up Anthropic just dropped a new model
Anthropic launched Fable 5.1 alongside Mythos 5.1 today. They use the same underlying model, but have different safeguards. Only Fable 5.1 is generally available.
High level stats: 1M token context window, 128K max output, June 2026 knowledge cutoff
Biggest improvements (according to Anthropic): long-running coding, scientific research, computer use, document creation, and complex workflows involving multiple tools
Improved Benchmark results (vs. Fable 5): Terminal-Bench-Science: 52.6% vs 24.7% Terminal-Bench 4.0: 55.8% vs 42.0% AutomationBench: 31.4% vs 17.1% GDPval-AA: 1,853 vs 1,723 CursorBench: 73.4% vs 70.5% OSWorld strict: 41.7% vs 36.1%
Worse Benchmark results: ARC-AGI-1: 97.5%, slightly below Fable 5 at 98.5% ARC-AGI-2: 90.0%, below Opus 5 at 90.42% and GPT-5.6 Sol at 92.5% SWE-bench Multimodal: 54.7%, below Opus 5 at 59.4% HealthBench Professional: 62.1%, slightly below Fable 5 at 63.3%
Base pricing is unchanged: Input: $10 per million tokens Output: $50 per million tokens 5-minute cache writes: $12.50 per million tokens 1-hour cache writes: $20 per million tokens The major change is cache reads: $1.00 → $0.25 per million tokens
And most excitingly, approximately 60% fewer cyber-safeguard interventions per Claude Code session → tweet link
@ivanfioravanti · 2026-09-01T11:26
Sherry quantization method by Tencent, combined with large Hy4 Preview is able to deliver incredible results! 1.5TB to 214GB using 1.25 bits weights while keeping incredible accuracy!
BF16 vs Sherry: - MCP Atlas 83.7→83.2 - SWE-Bench multi 82.9→81.3 - MRCR 81.3→81.1 - IFBench 73.5→72.5 → tweet link
@ivanfioravanti · 2026-09-01T18:37
My main driver for coding is now GLM 5.3 for nearly everything, Kimi K3 when I need design/frontend top capabilities.
I really don't feel the need for Fable 5.1
At least not for my current needs🤷🏻♂️ Do you really need Fable 5.1 level in your daily activities? → tweet link
@victormustar · 2026-08-31T19:08
RT @osanseviero: Today we're releasing TimesFM 3.0 on Hugging Face
- Open foundation model for time series forecasting
- Complex multivari… → tweet link
@jezell · 2026-09-01T06:25
RT @runwayml: Today, we're sharing new research on Solaris, our first Interface World Model.
Solaris is a new kind of operating system tha… → tweet link
Agentic Engineering & Tools
@Teknium · 2026-08-31T20:23
Hermes Agent v0.21.0 is now out!
- Bots Mode
- Agent 2 Agent Comms
- Persistent Multi-Gateway Connections
- Subagent Steering
- Expanded Connectors Access
and a lot more!
Check out the release notes below → tweet link
@kunchenguid · 2026-08-31T20:08
once upon a time i was working with ~10 agents in tmux. i got all the agents on full cylinders and it felt great. i walked away to make coffee
when i came back, the screen showed an empty terminal with nothing on it. it took me a few seconds to realize it was a tmux server crash. my "agent civilization" just got a team wipe
that was probably one of the most painful experiences in all my time working with agents. i had to manually figure out which agents were working, what their last state was, which worktrees were clean vs dirty, which PRs were raised vs not, how to get them to correctly resume
this team wipe led to me putting down a founding principle in firstmate - restart should be a non-event. it's hard to fully achieve that, but firstmate pushed pretty far
i can kill the firstmate session and crewmates would still keep working. restart firstmate and it would catch up on what the fleet had done
i can have dozens of second mates and crewmates running in the fleet doing whatever, and i can literally pull the power plug of my mac, restart it, and firstmate will reconcile everything and in a few minutes the entire fleet would be doing whatever it was doing earlier
i can have multiple remote machines and any one of them can die any time. whenever they are up, all the crew there would continue working
this was done by being disciplined about where to place what kinds of context
the agent's context window is like our computer's "memory". they can get lost, cleared, compacted, or screwed up in various ways. we should keep it intended for things that agent's next request likely needs direct access to
anything that needs to serve as a durable record should not be trusted to stay only in the context window. they must be pushed and persisted onto the disk, and made discoverable by the agent. this can be done through tools like beads, or external project management tools like jira/linear. firstmate does this mostly through a tool called "tasks-axi" which stores records in a markdown file locally
if you don't have a system for persistent record keeping yet, highly recommend looking into getting one set up. this saved my agent civilizations so many times → tweet link
@kunchenguid · 2026-09-01T02:12
everyone is running into a wall with agentic engineering right now
you don’t hear people talking about it because 1. they profit from selling you “solutions” 2. they want to look smarter than others
so let me be the whistleblower - the wall is called “how does it feel”. agents can’t do it
they can walk right past the ugliest UI or the most obvious bug and don’t say a thing unless that’s what you asked
they can take many screenshots and burn through my tokens but they can’t tell me if my landing page animation looks cool
they can click through my app but they won’t feel the dopamine hit when i physically drag my finger on the touch screen and feel the command dial flow with me in @theSSHHIP
this wall is structural. it’s actually load-bearing 😉 → tweet link
@KingBootoshi · 2026-08-31T20:15
WOW training a neural net is so delicate. It's like taking care of a VERY sensitive bonsai tree
I'm currently working on one of my first personal neural nets (NN)s I've made, which is speech recognition of my voice.
Siri has this, when you say 'hey siri' it has a small neural network running on your phone that recognizes your voice
Well I wanted one for my own voice agents and I wanted to see how to make one myself
I found a voice data sets online (like VoxCeleb) and I had a LOT of recordings of my own voice on my computer
The first model I trained actually worked instantly, against other voices vs mine, the NN was able to recognize my voice, and deny other peoples voices
but then I noticed that in complete silence, the neural net would almost recognize me. why ???
turns out, the data I trained on (that contained my voice) had a LOT of samples of complete ambient silence. so the NN was trained to ALSO recognize ambient silence as me
to combat this, i did two things
-
i cleaned my dataset to ONLY contain my voice, no silence. my agent successfully did this in a one shot prompt
-
i found (and recorded) more ambient noises, to train the neural net that SILENCE is NOT my voice
now it has a 100% pass rate, does not accept silence as me, and completely rejects other peoples voices
the actual CODE/parameters of training the neural net literally stayed identical. the only thing I did was clean up the data a bit, and BOOM perfect NN
now i kind of understand my role as a data labeler, because it seems to be NNs can quite literally learn ANYTHING given clean data (and that kind of freaks me out) → tweet link
@thdxr · 2026-09-01T16:42
one side effect of treating LLMs like humans is people keep using them ineffectively
you can't break a task up into 5 pieces and work on all of them in parallel and non-linearly
but the agent can. you don't even have to tell it to, if it has the right environment it will do it → tweet link
Developer Tools & Local AI
@ollama · 2026-09-01T05:49
Ollama’s Pro, Max, and Team plans now use transparent per-token pricing.
Based on your feedback, every plan includes a monthly pool of usage credits.
If you’re on an existing Pro, Max, or Team plan, your plan continues to work as-is. You can upgrade to the new pricing anytime in your Ollama account settings.
Every plan includes:
-
High-performance access to the latest open models, at published per-token rates
-
Works with popular coding agents, including Claude Code and Codex, plus an API for your own tools
-
Monthly usage credits included with every plan
-
Zero data retention, hosted in the US and Europe, plus Singapore for a limited set of Qwen models
-
No service fees or hidden limits
Pro: $20/month, includes $60 of monthly usage Max: $100/month, includes $300 of monthly usage Team: $500/month, includes $1,000 of shared monthly usage for unlimited users
The free plan now includes a small amount of monthly usage for a set of starter models.
Learn more directly from Ollama's pricing page: https://t.co/8GUwSqL48X → tweet link
@sudoingX · 2026-09-01T14:49
open weight ai fight film is here anon. two of the biggest open models that fit one dgx spark, the exact same 157 line octopus invaders game prompt, the same hermes agent harness and i recorded both builds end to end.
episode 4 of the local arena is the full head to head.
what this tape shows is ling 3.0 flash from @AntLingAGI builds like a craftsman with the spec pinned to the wall, 2,366 lines of js, every size and speed lifted literally from the prompt into config, 4 octopus types with pixel grid rendering, boss ai, damage floaters, screen shake, a game that feels finished the second you touch it, 18 of 18 tests green.
ling 3 flash also pays for that precision, decode opens at 39.5 tok/s and decays as the build deepens, 12.9 at 24K, 7.0 tok/s at 51K, 3.7 tok/s at 107K, and the full build stretched across two nights and 8 box deaths on the way.
laguna s 2.1 from @poolsideai builds like a machine on a deadline. same spec, all of it wired, in 1,888 lines, 20 percent leaner, and the whole file set landed in about 50 minutes of serve time. it never slows down because dflash keeps decode flat around 19 tok/s at every depth with bursts past 50 tok/s on clean code runs.
both models crashed my dgx spark. that is a gb10 serving story, not a model story, and both came back from disk and kept building, laguna even resumed its own crashed session and continued its own plan with no reprompt.
so the fight is precision that costs time against completeness at constant speed. i have every number, both games run, and the tape is below. you have everything i have now, call it. → tweet link
@sudoingX · 2026-09-01T06:28
to every AMD owner running qwen 3.8 27b dense or thinking about it, my repo holds 12 amd speed submissions now, every number a paired baseline vs flag a/b measured by the person who owns the card.
here you go, you might find yours:
rx 7900 xtx on windows vulkan: 41.0 → 85.4 tok/s, the amd record rx 7900 xtx on linux vulkan/radv: 28.8 → 70.7 rx 7900 xtx on linux rocm 10: 36.3 → 62.6, landed tonight rx 7900 xtx launch row: 30.7 → 43.9 rx 7900 gre 16gb: 28.7 → 47.8 radeon ai pro r9700 32gb: 27.0 → 43.3 2x rx 9070 vulkan: 22.1 → 41.6 rx 9060 xt 16gb: 15.2 → 28.7, nearly doubled strix halo, ryzen ai max 395: 11.9 → 28.7 vulkan, 11.5 → 23.7 windows, and a rocm rig holding 95%+ acceptance at draft depth 12 radeon 890m igpu: 2.7 → 5.7, a laptop igpu doubled its own speed
look at the xtx block again. one card, four software stacks, 43.9 to 85.4 with the same flag. on amd the stack you serve from is worth more than the silicon you paid for, and this table is the only place that spread is measured side by side.
and the table has holes with names on them. no 7900 xt, no 7800 xt, no single 9070 or 9070 xt, the entire rx 6000 generation missing, not one instinct card. running 27b dense on any amd metal, open a pr, your row becomes the reference
https://t.co/omwN7wRqUf → tweet link
@MengTo · 2026-09-01T13:29
I pretty much one-shotted this Japanese garden landing page with MiniMax Code.
It got the design and frontend right, then generated the ink-wash images, ambient video, and soundtrack inside the same project.
I usually need Codex or Claude Code plus Higgsfield and several MCPs for a page like this. MiniMax had everything built in, and it's at a fraction of the token cost. → tweet link
@victormustar · 2026-09-01T14:56
Hugging Face is (also) the best place to write about AI.
Your articles are now automatically linked from your model pages, so anyone looking at the model finds what you wrote about it. We shipped an improved UX to make writing, drafting and publishing better. We made it great to use as a team: add coauthors to work together and have everyone credited on the post.
Get HF Pro at $9/month or subscribe your org to Team or Enterprise to get started. → tweet link
Software Development & Open Source
@jezell · 2026-09-01T16:18
Been moving Flocker builds to Bazel for the past few days. Really is a really large build at this point, not so much because Flocker itself is massive, so much as that building ports of pyodide, flutter, wgpu, gpui, blitz, libreoffice, etc. and running their tests across a bunch of flocker supported platforms concurrently is a surefire way to run out of resources. Being able to serialize builds from concurrent agents in a queue instead of them all fighting for the same resources is one of the nice immediate benefits. → tweet link
@jezell · 2026-09-01T02:40
RT @JeremyCMorgan: EVE Online is finally leaving Stackless Python 2.7 after 16 years. 2.4 million lines through futurize, then manual revie… → tweet link
@Teknium · 2026-09-01T08:18
We just crossed our 100,000th PR from contributors and team member contributions
Crazy lol → tweet link
@LinusEkenstam · 2026-09-01T09:17
RT @okuiux: Just launched 🚀 Scratched after an all nighter!
Probably the fastest video editor ever built ;)
Try it now [👇 link below]
So… → tweet link
Hardware & Tech Industry
@TrungTPhan · 2026-09-01T17:43
John Ternus updated his X before he updated his LinkedIn. Very bullish for Apple. https://t.co/9RBY9VrAXu → tweet link
@ivanfioravanti · 2026-09-01T16:36
I just invested in Apple shares! I bet John will push hardware and AI to the next level 🚀 A new era has begun! → tweet link
@MilksandMatcha · 2026-09-01T17:44
It is interesting to see the industry move from 12Hi back toward 8Hi HBM just as DRAM stacking is accelerating.
Going vertical does not make the fundamental constraints of memory disappear. As stacks get taller, thermals, power, yield, packaging, and reliability become increasingly difficult architectural problems.
HBM stacks DRAM vertically beside compute. The next frontier (as seen at hot chips) is stacking DRAM directly on compute, turning the interface from the edge of a chip into its entire surface area. → tweet link
@TrungTPhan · 2026-09-01T15:28
James Dyson is 79 and still finding ways to apply hardcore engineering to random consumer products.
He was in fine form pitching Dyson’s new $499 CameraJet toothbrush:
“The first toothbrush with a camera. The first toothbrush with a jet. The first toothbrush to find the space between the teeth and floss while you brush.”
Dyson says it created an ML algorithm trained on 470,000 mouth images to recognize the gaps to clean with water flossing. → tweet link
@FrameworkPuter · 2026-08-31T23:47
We've come a very long way... Framework Laptop 13 Pro is on @PCMag's Best Battery Life list. They hit 27 hours in their battery life test. https://t.co/rp7RfSbOGu → tweet link
@ivanfioravanti · 2026-08-31T22:09
Zai is at 2B$ ARR multiplying last week recurring revenues by 52.
Let that sink in! → tweet link