AI and IT News Recap: September 18, 2026: Anthropic Says Claude Now Leads a Quarter of Its Own R&D, a 9.8 Check Point Flaw Runs Code as Root, and Gyazo Loses 23 Million Accounts
By Noah Smith, Owner & Consultant, KeyChange Technologies · September 18, 2026

A short day for model launches and a heavy one for patching. Here is the AI and IT news that actually mattered in the last 24 hours.
📌 The AI and IT news at a glance
- Anthropic put a number on AI building AI. Claude now "leads" 26% of the company's own model research and development, up from under 1% in February.
- OpenAI launched Astra for Law. GPT-6 Astra, configured with a legal search index, shipping to selected firms with 26 partner plugins.
- Anthropic opened the Life Sciences Verification Program. Vetted biology teams get models with safeguards loosened for scientific work.
- Check Point patched a 9.8 flaw in its management servers. An unauthenticated attacker could run code as root on the box that controls firewall policy.
- Gyazo lost 23.62 million user records. Plus roughly 490 million image metadata records, including the IDs in image links.
- DNS had a genuinely bad day. BIND 9 fixed fourteen flaws and Unbound fixed a critical remote code execution bug, on the same day.
- A China-aligned espionage group rolled out a new backdoor. FamousSparrow has been running SparroWocky across Latin America since August 2025.
- Docker patched a sandbox escape aimed squarely at AI coding agents. Malicious code inside a sandbox could read and change files anywhere on a Mac.
🔝 Top story: Anthropic says Claude now leads a quarter of its own R&D
Anthropic published a set of measurements on Wednesday that it wants other frontier labs to adopt, and the headline figure is striking. Using an automation scale developed by Epoch AI that runs from AL0 (no AI involvement) to AL5 (fully autonomous), the company built what it calls the Anthropic R&D Automation Index by cataloguing roughly 15,000 granular model R&D tasks pulled from internal work records, organizing them into a 542-node tree, and having a Claude judge rate how automated each branch is. As of August 2026, Claude "leads" 26% of Anthropic's AI research and development work, meaning it completes most of a given task end to end from a high-level prompt while a human supervises. The share of work at or above "AI collaborates" is above 90%. Anthropic says Claude is not operating fully autonomously for any measured subset of that work. In February 2026, the "leads" figure was under 1%.
The post covers two other measurements. Anthropic says roughly 30,000 agents were doing research and engineering work on its most-used internal platform at any one time in August, that 100% of their actions pass through a real-time monitor before execution, and that of more than a billion agent decisions analyzed over the month, 0.002% (about 1 in 47,000) were blocked. An offline monitor flags around 100,000 transcripts a week, of which roughly 50 escalate to human review. On compute, a snapshot of the week of July 13 to 20 showed about 6% of AI R&D compute went to safety work, and about 12% of the compute spent on AI-driven AI R&D did. The company says it plans to embed independent third-party evaluators with access comparable to its internal risk teams.
In short: Anthropic published three new transparency metrics and disclosed that Claude now leads 26% of its own model R&D, up from under 1% in February 2026.
What it means for your business: Nothing changes in your tooling today, but the pace of capability improvement in the models you are buying is now partly a function of those models improving themselves, which is a reasonable input into how long you plan around any single vendor or version.
My take: The number everyone will quote is 26%, and the number that deserves more attention is the jump from under 1% in February. Also worth reading carefully: Anthropic used its own models to grade its own models, which it says up front, and the human spot-checks agreed with the judge about as often as humans agreed with each other. That is honest about the limitation rather than hiding it. Whether other labs publish the same numbers is the real test of whether this becomes a standard or stays a press release.
🤖 AI
OpenAI launches Astra for Law
OpenAI introduced Astra for Law, which pairs GPT-6 Astra with a legal search index and custom instructions for legal analysis and writing. The index covers U.S. case law, statutes, regulations, court rules and administrative decisions across more than 230 million URLs, with case-law coverage coming through a collaboration with the Free Law Project, the nonprofit behind CourtListener. On 200 questions from the private validation set of Vals AI's Legal Research Bench, OpenAI reports Astra for Law passed the overall correctness check on 54.0% of questions versus 38.7% for GPT-6 Astra using web search alone, a 40% relative improvement, with both systems at their highest reasoning effort.
Access is initially limited to selected law firms through a Trusted Access program in ChatGPT and Codex, appearing in the model picker as "GPT-6 Astra Law," with API availability described as coming soon. Eligible firms get Zero Data Retention on the API, and ChatGPT Enterprise usage is excluded from human review by default. OpenAI also launched 26 partner-built plugins connecting ChatGPT to tools like iManage, Intapp, Relativity, Clio and Thomson Reuters HighQ, plus nine community plugins carrying 47 custom skills. Separately in the same announcement, ChatGPT for Word became generally available.
In short: OpenAI released Astra for Law, a version of GPT-6 Astra configured for legal work with its own case-law search index, available first to selected firms.
What it means for your business: If you are not a law firm, the interesting part is the pattern: a frontier model wrapped in a domain-specific search index and instructions, sold as an industry product. Expect the same shape to arrive in accounting, healthcare and construction, and expect your vendors to start pitching it.
My take: The benchmark jump from 38.7% to 54.0% is real but it is also a reminder that the ceiling on legal research is nowhere near solved. Nearly half the questions still fail the correctness check at the highest reasoning setting. That is a useful research assistant and it is not a lawyer. The governance work with Latham & Watkins on ethical walls and information permissions is arguably the more consequential piece here, because that is what has actually been blocking firms, not model quality.
Source: Introducing Astra for Law, OpenAI, September 17, 2026
Anthropic opens the Life Sciences Verification Program
Anthropic opened applications for its Life Sciences Verification Program, which gives vetted life science organizations access to Mythos, Opus and Sonnet models with classifiers tuned to be more permissive for biology work that the generally available Fable models currently block. Applicants go through a review of research credentials, security standards and ethical research oversight, then apply for one of two grant types. Standard Use grants cover most research and development workflows, extend to whole teams, and renew annually. High-risk Use grants remove all life sciences safeguards, apply to a single research project rather than a team, and renew every six months. High-risk grants are available today for Opus 5 and Sonnet 5; for Mythos, Anthropic says it is working with the US government and access remains limited to a small set of additionally vetted entities.
The mechanics are notable. For this traffic, Anthropic is shifting from real-time blocking to offline monitoring against each organization's stated use cases, which requires 30-day data retention on flagged activity. The company says that data is compartmentalized, cannot be used for training, and cannot be accessed by its own life sciences research teams. Dozens of organizations were onboarded in early access, including Xaira Therapeutics, Edison Scientific and Manifold Bio, and Anthropic expects to enroll hundreds in the first week. The beta is not available to BAA-enabled organizations, so customers handling PHI need separate non-BAA orgs.
In short: Anthropic launched a verification program that gives credential-checked life science teams access to models with biology safeguards relaxed, backed by offline monitoring and 30-day retention on flagged activity.
What it means for your business: This is a template for how regulated industries will get access to capabilities that are blocked by default: prove who you are, state what you are doing, and accept monitoring in exchange. If your sector is currently frustrated by over-blocking, this is the shape of the answer.
My take: Trading real-time blocking for after-the-fact monitoring is the honest trade-off, and Anthropic names it plainly rather than dressing it up. It also means the safety story now depends on the customer's own admins responding to flags within agreed timeframes, which is a meaningful shift of responsibility onto the buyer. Worth reading the retention terms closely before anyone in your org signs up.
Source: Introducing the Life Sciences Verification Program, Anthropic, September 17, 2026
🛡️ IT and security
A 9.8 Check Point flaw lets unauthenticated attackers run code as root
Check Point disclosed a critical stack overflow in the login process of its Security Management and Log Servers, tracked as CVE-2026-91843 and rated 9.8 by the vendor. Because the flaw sits in code that handles requests before a user authenticates, an attacker with no credentials can reach it over the network and run code as root on the server that controls firewall policy and administrator access. Internet scanning firm Censys said the overflow is triggered by a login request carrying a very long username, and that it saw no public proof-of-concept exploit as of September 16. Check Point says it has no indication of exploitation, and CISA recorded exploitation as "none" on the CVE record on September 17.
Fixes ship through Check Point's LivePatch channel under advisory sk1000155. Affected branches include R82.10 at Jumbo Hotfix Take 44 or below, R82 at Take 126 or below, R81.20 at Take 166 or below, and R81.10 at Take 190 or below, along with R81, R80.40, R80.30, R80.20, R80.10 and R80, all of which are end of support and get no fix. Censys and an NHS England Digital alert also list R82.20 as affected even though Check Point's CVE record does not. If you have automatic updates on, verify rather than assume: the cplp list command shows installed LivePatches, and when Check Point pushed fixes for two VPN certificate flaws last week, customers reported the package had not reached them on announcement day. This is the fifth critical flaw since July 22 reachable on a Check Point management server without logging in.
In short: Check Point patched CVE-2026-91843, a 9.8-rated pre-authentication stack overflow that lets an unauthenticated attacker run code as root on Security Management and Log Servers.
What it means for your business: If your network gear is Check Point, ask whoever manages it to confirm the LivePatch is actually installed today, and confirm that management access is not reachable from the internet. If you are on an end-of-support branch, the only fix is an upgrade.
My take: Five unauthenticated-reachable critical flaws in the management plane since late July is a pattern, not a run of bad luck, and it is fair to ask Check Point about the state of that codebase. For everyone else the lesson is duller and more useful: the management console is the crown jewel, not the firewall itself, and far too many of them are still reachable from places they have no business being reachable from.
Gyazo breach exposes 23.62 million user records
Helpfeel, the Kyoto-based company behind the image-sharing service Gyazo, disclosed a breach affecting about 23.62 million user records, including email addresses and password hashes. The incident also exposed roughly 490 million image metadata records, mostly covering images from January 2019 or earlier, and those records include the IDs that make up Gyazo image links. Helpfeel said those IDs could be used to view images without permission and that it has temporarily disabled viewing of some of them.
The attacker got in through a vulnerability in Gyazo's image upload server, ran arbitrary commands on Helpfeel's systems, and reached the Gyazo database. The company has not said what kind of flaw it was. It is asking every user to change their Gyazo password and to change it anywhere else they used the same or a similar one, and to watch for suspicious messages referencing the incident.
In short: Gyazo's operator disclosed a breach exposing about 23.62 million user records and roughly 490 million image metadata records, including the IDs embedded in shareable image links.
What it means for your business: Gyazo is a screenshot tool, and people screenshot things like dashboards, invoices, internal chats and configuration screens. If anyone on your team used it, the exposure is not just the account, it is whatever was in those pictures. Password reuse is the other immediate risk.
My take: The 490 million metadata records are the part worth sitting with. Link IDs mean the images themselves may be reachable without a login, and screenshots are exactly the kind of casual, unclassified data that nobody inventories. This is a good prompt to ask what small free tools your team quietly relies on, because that list is almost always longer than anyone expects.
Two DNS servers, one bad day
The Internet Systems Consortium released BIND 9.20.29 and 9.21.26 to fix fourteen security flaws disclosed on September 16. One affects any BIND server answering DNS-over-HTTPS: a sender with no credentials can crash the named process with a single request carrying an invalid SIG(0) signature, if the sender closes the connection before the signature check finishes. BIND 9.20.29 fixes all fourteen; 9.21.26 fixes thirteen, because CVE-2026-19662 does not affect the 9.21 branch. ISC says it is not aware of any of the fourteen being exploited and lists no workarounds.
The same day, NLnet Labs disclosed that every release of the Unbound resolver before 1.26.1 carries a critical heap overflow in its DNSSEC validator, tracked as CVE-2026-81642. An attacker who controls a malicious zone and queries a vulnerable resolver can trigger it, enabling remote code execution. Unbound 1.26.1 fixes it along with eight other flaws, one of which, CVE-2026-82717, is a heap corruption bug in CNAME synthesis reported by Ben Morris of Anthropic that could also lead to code execution under certain systems and compilation options. NLnet Labs has not reported exploitation of either, and CISA marked CVE-2026-81642 exploitation as "none."
In short: BIND 9 and Unbound, the two most widely deployed open-source DNS resolvers, both shipped critical security updates on September 16 and 17.
What it means for your business: Most small businesses do not run their own resolver, so the real action item is a question for your ISP, hosting provider or managed IT firm: are your recursive resolvers patched? DNS failures do not look like a security incident, they look like the whole internet being broken.
My take: Two independent DNS codebases patching critical bugs in the same 48 hours is coincidence, not a coordinated event, but the effect on defenders is the same. Neither is reported as exploited, which is the good news, and neither has a workaround, which is the bad news. Patch is the only lever. One small bright spot: an Anthropic researcher reported one of the Unbound bugs, which is the kind of AI-lab-to-open-source contribution that is quietly becoming normal.
FamousSparrow deploys a new backdoor across Latin America
ESET reported that the China-aligned state-sponsored group FamousSparrow has been deploying a previously unreported modular C++ backdoor called SparroWocky against targets in multiple Latin American countries since at least August 2025. Researchers Alexandre Côté Cyr and Romain Dumont describe an architecture and anti-analysis tooling that indicates strong knowledge of Windows internals. The malware gets its name from early versions containing the first stanza of Lewis Carroll's nonsense poem Jabberwocky.
The group, which shares some overlap with the clusters tracked as Earth Estries and Salt Typhoon, appears to have replaced its older SparrowDoor implant with this one. This is long-running espionage tradecraft rather than opportunistic crime, and the target set is governments and larger enterprises rather than small businesses.
In short: ESET attributed a new modular backdoor called SparroWocky to China-aligned FamousSparrow, used against Latin American targets since at least August 2025.
What it means for your business: Direct relevance is low unless you operate in the region or supply organizations that do. The indirect relevance is supply chain: state groups routinely reach large targets through smaller vendors, and "we are too small to be interesting" has not been true for a while.
My take: The detail I keep coming back to is that this has been running since August 2025 and is only being reported now. Thirteen months of dwell time on a campaign run by a group that researchers already track closely is a fair measure of how hard this category of attack is to see. Nothing to patch here, but it is a useful calibration on what "detected quickly" actually means.
🧰 New tooling for builders and everyday AI use
Docker patches a sandbox escape aimed straight at AI coding agents
Docker disclosed that malicious code running inside a Docker Sandboxes virtual machine on macOS could escape the shared project directory and read or change files anywhere else on the host, running with the rights of the host account that runs the VM. The flaw, CVE-2026-77179, is rated Critical at CVSS 9.4, affects versions 0.28.0 up to but not including 0.42.0, and was fixed in 0.42.0 on September 7, with the advisory published September 15. The escape went through the virtio-fs host server, which followed symlinks when reopening a removed file from a stored path, letting a guest replace a parent directory with a symlink and reach outside the workspace. The same release fixes CVE-2026-79994, rated High at CVSS 8.7, in the relay that lets a sandbox connect to Unix domain sockets inside its workspace.
The context matters. Docker Sandboxes exists specifically to run each AI coding agent in its own small virtual machine with your project folder shared in, and Docker's own documentation says the hypervisor boundary "is the isolation control, not in-VM privilege separation." In other words, the thing that broke is the exact thing you installed it for. Docker has reported no exploitation, and neither CVE is in CISA's Known Exploited Vulnerabilities catalog. Update to 0.42.0 or later; as of September 17 the newest release was 0.43.0. If you cannot update, Docker advises using clone mode and avoiding read-write host mounts, though clone mode only works on Git repositories and protects the repo from changes, not from being read. Untracked files like .env stay readable inside the sandbox either way.
In short: Docker patched a critical flaw that let code inside an AI coding agent's sandbox read and modify files anywhere on the macOS host, fixed in Docker Sandboxes 0.42.0.
What it means for your business: If anyone at your company runs AI coding agents on a Mac inside Docker Sandboxes, update today. More broadly, the safety of "just let the agent run in a sandbox" depends entirely on that sandbox holding, and sandboxes are software with bugs like everything else.
My take: This is the most practically useful story of the day for anyone building with AI agents. The whole premise of handing an agent broad permissions is that the blast radius is contained, so a container escape is not just another CVE, it invalidates the assumption the workflow rests on. Credit where due: no exploitation reported, a fix shipped before the advisory, and named credit to the researchers. But if you are running agents with
sudoinside a box you trust, put "keep that box patched" on the same list as your servers.
Missed yesterday? Catch up with the September 17, 2026 AI and IT news recap.