OpenAI has published a case study describing how Asana used GPT-6 Astra in Codex to optimise the browser agent in its StackAI automation platform. The resulting workflow, running on OpenAI's GPT-6.1 Sol model, averaged an estimated $0.47 in model costs and about four minutes per run. That is 76x cheaper and 5x faster than the original production setup, which cost at least $36.21 per run and took at least 22.5 minutes.
According to the study, GPT-6 Astra identified a key inefficiency: the agent cached its fixed instructions and tool definitions, but not the growing history of page text and screenshots it gathered, so every request resent that history at full price. The agent also dropped older screenshots and trimmed text at nearly every step, which altered the history and prevented caching. Asana's StackAI CTO Frank Hidalgo selected three fixes to test: extending caching to the browsing history, increasing the amount of text retained, and removing screenshots in batches rather than at every step.
The best-performing policy allowed screenshots to accumulate to 20 before cutting back to the most recent one. In a 144-run study comparing GPT-6.1 Sol with three other frontier models, this policy combined with a larger history budget produced the lowest costs. On GPT-6.1 Sol alone, the new caching and screenshot policy cut cost 4x, from $1.97 to $0.47 per run, with 89% of input coming from cache at 5% of the uncached price. The larger history budget also improved reliability: all 18 runs with the larger budget produced correct answers, versus 3 of 18 with the smaller budget.
Hidalgo estimated the investigation would have taken one to two months by hand; with GPT-6 Astra in Codex, it took about a week. Asana CPO Arnab Bose said the work demonstrates what teams of humans and agents look like in practice, with an engineer setting direction, GPT-6 Astra running experiments, and the results going through Asana's Command platform to production.