Anthropic engineers recently made claude.ai and the Claude desktop app about three times faster during a two-week sprint in August, according to a Help Net Security report. The team merged more than 3,000 changes, with Claude itself identifying bottlenecks and writing fixes, while humans set goals and approved every change. Notably, none of the changes caused a customer-facing incident or rollback, a strong reliability record for such a large refactor.
The sprint used an internal research model comparable to Opus 5.5, running through the Claude Tag beta. Claude built benchmarks, opened pull requests, and monitored deployments, with over 150 threads active at once. To avoid noisy wall-clock timing, Claude switched to deterministic counts like CPU instructions—cutting instructions by 48% on one hot path led to a 78% real-time improvement. A CI check then prevented regressions by failing any change that increased the instruction count.
Some slowdowns were surprising. Highlighting a finished code block could freeze the page for about a second because em dashes and other non-Latin-1 characters forced V8 to store replies as UTF-16, pushing syntax-highlighting regexes onto a slower path. A 20-line fix resolved it. The team used nearly 200 short-lived feature flags, rolling the riskiest changes out to employees first and removing more than half by the end. Engineer Issac G. said, "You could not have convinced me this was possible even six months ago."
While the article focuses on performance, the security and privacy angle is the absence of incidents during a massive automated change. The disciplined approach—human oversight, feature flags, and regression checks—offers a template for safely applying AI-generated code to production systems without compromising user trust or system integrity.