Technology · 25 September 2026 · Evening edition
Microsoft builds on OpenClaw as developers push more AI work onto local machines
Microsoft has shipped a product built on OpenClaw, according to the open-source project’s creator Peter Steinberger, who said the companies had worked together since March to prepare the codebase for large-scale deployments. At the same time, developers were reporting faster local-model performance on consumer GPUs and Apple hardware, while Cognition announced a billion-dollar annualised revenue run rate. Together, the posts show commercial expansion alongside an active effort to make more AI work possible outside hosted services.
A commercial milestone, with the product still unnamed
Steinberger described Microsoft as a partner and open-source contributor, saying it had released a compelling product on top of OpenClaw. His announcement did not name the product, set out its pricing or provide deployment figures. The concrete news is his account of the collaboration and shipment, rather than a complete product launch description. Peter Steinberger
Cognition’s financial announcement, reshared by swyx, said the company had crossed $1 billion in annualised revenue run rate. That measure is not the same as revenue already earned over a full year, and the brief announcement says nothing about profitability. Separately, a post from The Information reshared by jezell reported that TypeSafe AI was talking to investors about raising $1 billion or more. Those were fundraising talks, not a completed financing round. Cognition announcement · Funding report
Old graphics cards, new software gains
Local-model developer sudoingX reported that restoring a multi-token-prediction head and applying a kernel fix raised Bonsai 2 27B output on an RTX 3060 from 26 to 50 tokens per second at fresh context. The author described the model as a ternary compression of Qwen 3.8 27B and said throughput fell to 18.3 tokens per second with 77,000 tokens already in the context window. The contrast matters: a headline speed at the start of a conversation is not sustained performance throughout a long task. Bonsai measurements
The same developer reported running Qwen 3.8 Flash Next in FP8 across two DGX Spark machines at 45 tokens per second with fresh context. GPU power during generation was put at 30–35 watts per device, with higher peaks. These are the author’s measurements of the GPUs, not a whole-system electricity figure or an independently reproduced efficiency comparison. Dual-Spark results
Apple hardware supplied another set of results. Ivan Fioravanti reported MTPLX 2.12 tests on an M3 Ultra with 512GB of memory, using Speed and Quality versions of Qwen 3.8 Flash Next. Reported peak decoding ranged from 61.6 to 96 tokens per second across prose and code tasks. The figures describe different test cases, not a single sustained rate. Fioravanti’s tests
A comparison posted by TheAhmadOsman put two DGX Sparks ahead of an M5 Ultra in DeepSeek V4 Flash prompt processing by a factor of 2.41, while the Mac led token generation by 6.4%. The author explicitly noted that the Mac lacked optimised kernels. Fioravanti likewise urged readers to wait for further optimisation before drawing conclusions about M5 Ultra performance. These qualifications make a simple hardware league table premature. Hardware comparison · Optimisation caveat
Agents gain browser, speech and automation tools
Several announcements concerned the software surrounding models. Posts reshared by jezell described Jev support for DSPy, including an optimiser for output thresholds, and a demonstration in which Jev made decisions while Lightpanda performed browser clicks to search for hostels. Both show specific integration work; neither short excerpt supplies a broad reliability evaluation. DSPy integration · Browser demonstration
Photon was said to have added streaming speech output with Qwen3-TTS models and Kokoro-82M, in a developer announcement reshared by alexinexxx. The excerpt cuts off before a complete first-audio latency claim. In another announcement, reshared by sqs, Amp’s installed Mac app was said to be able to spawn runners automatically. Photon speech · Amp runners
Beyond agent interfaces, Wasmer announced a PostgreSQL 18.4 server running on iOS, following its work with Python, Node.js and FFmpeg, according to a repost by jezell. The preserved excerpt establishes the announcement but not the server’s compatibility limits or suitability for production use. Wasmer announcement
Creative demonstrations meet the question of control
Meng To described a 47-minute tutorial using Opus 5.5, Mobbin MCP references and three.js to produce a landing page, brand guidance and advertising assets. The account is a practitioner’s demonstration of a design workflow, not a controlled comparison of models. Design tutorial
Developers also described different ways of supervising the work. Mario Zechner said he reviewed agent-produced changes as diffs in VS Code rather than watching individual editing calls. Kunchenguid outlined a personal model-routing system that separates interactive coordination, planning, implementation and difficult escalations. Both accounts place the organisation and review of tasks alongside model capability—not just the speed at which code appears. Reviewing changes · Task routing