Claude Opus 5 tops ARC & Cyber risk in open models - Hacker News (Jul 25, 2026)
Claude Opus 5 leads tough AI tests, Android ADB may tighten, Wasmtime upgrades WebAssembly, and a tiny Playdate proves 3D is possible.
Our Sponsors
Today's Hacker News Topics
-
Claude Opus 5 tops ARC
— Anthropic launched Claude Opus 5, a stronger flagship AI model for coding, automation, and knowledge work. Its lead on ARC-AGI-3 also shows that adaptive reasoning remains difficult even for frontier models. -
Cyber risk in open models
— A UK AISI and CAISI assessment says Moonshot AI's Kimi K3 is now the strongest open-weight cyber model they have tested, but it still trails top U.S. systems. The report highlights growing security concerns around accessible AI capability. -
Android ADB loopback under threat
— An Android IssueTracker discussion has raised fears that on-device ADB over loopback could be restricted to Wi-Fi only. That would affect Shizuku, Termux, accessibility tools, and other power-user debugging workflows. -
Wasmtime turns on Wasm GC
— Wasmtime 47 now enables WebAssembly GC and exception handling by default. This is a major step for running higher-level languages on Wasm with smaller binaries and more natural runtime behavior. -
Playdate gets real-time 3D
— A developer built a playable 3D software renderer for the Playdate handheld despite its tiny 1-bit screen and limited hardware. It is a strong example of optimization, smart tradeoffs, and creative game engineering. -
Hannah Fry wins math prize
— Cambridge professor Hannah Fry won the Leelavati Prize for public understanding of mathematics. The award recognizes her work making math accessible through broadcasting, writing, and digital media. -
Tokyo museum saves dead media
— The Extinct Media Museum Tokyo is preserving obsolete devices and recording formats through a hands-on collection. Its open approach helps document media history for researchers, creators, and the public. - 08
Apartment aquaponics keeps improving
— A small apartment aquaponics project in New York has matured into a stable, compact food-growing system. The update offers practical lessons on urban sustainability, fish care, and low-space gardening.
Sources & Hacker News References
- → Android May Restrict On-Device ADB and Break Shizuku Workflows
- → Anthropic Releases Claude Opus 5
- → Hannah Fry Wins International Leelavati Prize for Math Outreach
- → Developer Builds a 3D Renderer for the Playdate Handheld
- → ARC Prize Leaderboard Shows Strong Models Still Struggle on ARC-AGI-3
- → Kyber Seeks Head of Engineering to Scale AI Document Platform
- → How an NYC Apartment Aquaponics System Was Improved and Stabilized
- → Wasmtime 47 Enables WebAssembly GC and Exceptions by Default
- → UK and U.S. Agencies Assess Kimi K3’s Cyber Capabilities
- → Extinct Media Museum Tokyo Publishes Official Overview
Full Episode Transcript: Claude Opus 5 tops ARC & Cyber risk in open models
Even the latest flagship AI model can top a major benchmark and still score only about 30 percent on its hardest test. Welcome to The Automated Daily, hacker news edition. The podcast created by generative AI. It is July 25th, 2026, and I’m TrendTeller. Today: AI gets stronger but still hits real limits, Android developers worry about a possible ADB change, WebAssembly reaches an important milestone, and we’ve also got a few smart, lighter stories from math, retro media, and urban tinkering.
Claude Opus 5 tops ARC
We’ll start with AI. Anthropic has released Claude Opus 5, its new flagship model, and the big pitch is not just raw capability but better efficiency and stronger behavior on long, messy tasks. The company says it is better at coding, debugging, automation, and scientific work, and more reliable when it needs to check its own output instead of giving up early. What makes the timing interesting is that the ARC Prize leaderboard also updated, and Opus 5 currently leads the verified ARC-AGI-3 results. Even so, that top score is only a bit above 30 percent. So the story here is mixed in a useful way: AI systems are clearly improving, but the harder tests still show how far they are from flexible, human-like adaptation.
Cyber risk in open models
Staying with AI, there is also a new warning sign in cybersecurity. A preliminary assessment from UK AISI and CAISI looked at Moonshot AI’s Kimi K3 and found that it is now the most capable open-weight cyber model they have tested so far. It still falls short of the leading U.S. frontier models, but the gap is narrowing enough to matter. In practical testing, it made real progress on exploit tasks and simulated network attacks, even if it was not yet top tier. Why this matters is simple: when stronger cyber capability becomes available in more open models, the conversation shifts from hypothetical risk to operational risk.
Android ADB loopback under threat
On the Android side, a blog post is drawing attention to a possible future restriction on ADB connections made directly on-device. This is not an official Google announcement, but rather a concern based on an ongoing IssueTracker discussion tied to a security flaw. The worry is that ADB could end up limited to the main Wi-Fi interface, which would break loopback-based workflows used by tools like Shizuku, Termux setups, and other local debugging methods. The author’s point is that these are not fringe abuse cases. For many developers and power users, they are legitimate ways to automate tasks, support accessibility, and work around device limits. So the broader issue is not just security, but whether Android can tighten defaults without shutting down a small but important ecosystem.
Wasmtime turns on Wasm GC
Next, a meaningful milestone for WebAssembly. Wasmtime 47 now enables WebAssembly garbage collection and exception handling by default. That may sound like plumbing, but it is a big deal for the future of higher-level languages on Wasm. In plain terms, it means languages that rely on managed memory and normal exception behavior can target WebAssembly more naturally, without dragging along as much custom runtime baggage. That should help shrink binaries, improve efficiency, and make cross-language components easier to build over time. It is one of those infrastructure updates that most users will never see directly, but a lot of developers will eventually benefit from.
Playdate gets real-time 3D
One of the most fun engineering stories today comes from the Playdate handheld. A developer has built a real-time 3D software renderer for the tiny 1-bit device, starting from a simple question: is this little machine fast enough to do usable 3D at all? The answer turned out to be yes, but only with careful compromises. After experimenting with textures and visibility tricks, the project settled into a cleaner cel-shaded look that suits the screen much better than noisy full-detail scenes. What makes this worth noticing is not just nostalgia or novelty. It is a reminder that constraints can still produce inventive software design, especially when the hardware looks almost laughably limited on paper.
Hannah Fry wins math prize
And finally, a few lighter but worthwhile notes. Cambridge mathematician Hannah Fry has won the Leelavati Prize, one of the major honors for public communication of mathematics. It is recognition not just for academic work, but for making math feel relevant, understandable, and even entertaining to wider audiences. In Tokyo, the Extinct Media Museum is building a hands-on archive of obsolete devices and formats, from cameras and typewriters to old mobile tech, with an unusually open attitude toward photography and scanning. That makes preservation more participatory, which is refreshing. And in New York, a long-running apartment aquaponics project has been updated with practical lessons on making a compact indoor food system more stable and productive. Different stories, but they all point to the same idea: technical culture lasts longer when people share it clearly, preserve it carefully, and keep experimenting at human scale.
That’s it for today. Links to all the stories we covered can be found in the episode notes. Thanks for listening to The Automated Daily, hacker news edition, and I’ll be back tomorrow.
More from Hacker News
- July 23, 2026 Google AI adoption and EU scrutiny & Cruller trims Bun for production
- July 22, 2026 OpenAI models cause real breach & Anthropic settles book piracy claims
- July 21, 2026 AI challenges old math problems & Chinese models reshape AI economics
- July 20, 2026 AI finds WordPress exploit chain & DIY retrofit revives bowling tech
- July 19, 2026 Local speech AI shrinks & Open models challenge incumbents