Every step a browser proxy takes may result in the re-sending of page text, screenshots, tool definitions, and previous operations to the model. Intuitively, one might think that saving money would mean switching to a cheaper model or deleting old screenshots as soon as possible. However, cases published on October 8th and 9th by Asana and OpenAI demonstrate that this intuition is not always correct. After the StackAI team under Asana changed the way proxy history was cached, the rhythm of screenshot cropping, and the context budget, the optimized GPT-6.1 Sol process had an average cost of about $0.47 per model run and took approximately 4 minutes in a controlled test. Compared to the original production configuration using another model at the same price point, the estimated cost was reduced by 76 times, and the speed was about 5 times faster. The multiples mentioned in the title are striking, but they represent the results of an experiment for a specific task and a specific control group, and are not a general promise of cost reduction for all browser proxies.
This case is different from simply showing “AI helping engineers write code.” StackAI CTO Frank Hidalgo handed over the problem to GPT-6 and Astra, who then studied it within Codex. They first sorted out how the requests were constructed, then recorded the cost of each step, and made comparisons between different caching strategies and historical budgets. People were responsible for setting goals, selecting solutions, and reviewing changes, while the models undertook a large amount of experimentation and analysis. Asana said that what was originally estimated to take one to two months to study was completed in about a week; however, what is truly worth reusing by peers is not the project timeline slogan, but rather the experimental design and the records of each step left behind.
Cache invalidation; the issue does not lie with the cache switch itself.
The original proxy had cached fixed prompts and tool definitions, but not the growing history of web pages. Even more troublesome was that it deleted and cropped old screenshots and text with almost every request. Many model services only reused the longest unchanged prefixes for prompt caches; once the earlier history was modified, subsequent large sections of content also lost the opportunity to be reused. In this way, although deleting content made each request shorter, it could result in the same information having to be entered repeatedly at the original cost. The proxy might also need to re-access the page due to the deletion of key facts, increasing the number of steps and waiting time. Therefore, the focus of optimization shifted from "as short as possible" to "as stable as possible, with batch deletions and modifications only when necessary."
There are three modifications tested by the team: overwriting the cache with the browsing history; expanding the history budget from 120,000 characters to 480,000 characters; and only cropping screenshots in batches after a certain number have been accumulated. Under optimal conditions, a maximum of 20 screenshots are retained, with the latest one being kept. The combination of these methods is not about using as much data as possible, but rather about ensuring that a longer section of the request remains unchanged across consecutive steps, which facilitates repeated caching. The experiment published by Asana covered four models, two levels of budget, six types of caching, and six screenshot strategies. Each condition was repeated three times, for a total of 144 runs, with an additional 12 supplementary tests. The task was fixed to collect six pieces of information for each of 32 books from the publicly available book catalog. This approach allows for a comparison of cost, speed, and accuracy, but it cannot automatically represent all websites, login scenarios, or long-duration tasks.
The data also reveals the significance of “model selection” and “process optimization.” For the Model B used in the original production configuration, the optimization process reduced the cost per estimation from at least $36.21 to $1.24, a reduction of about 29 times. When switching to the optimized configurations of GPT-6.1 and Sol, the cost was further reduced to $0.47; within the same budget for Sol, changing the caching and screenshoting strategies reduced the cost from $1.97 to $0.47 as well, a reduction of about four times. In the original configuration, some tasks were not completed before reaching the maximum number of steps, so “at least $36.21” represents the lower bound, and the comparison of a 76-fold reduction is also influenced by this baseline setting. If all these improvements are attributed to a single new model, it would overshadow the main engineering findings.
Completion rate is equally important. Asana states that with a smaller historical budget, only three out of 18 runs of GPT-6.1 Sol produced an answer; however, after increasing the budget, correct answers were given in all 18 runs. There are 192 factual points to collect for this task, and if the agent forgets the page content too early, it may not be possible to complete even with each step being relatively inexpensive. After optimization, about 89% of the inputs are read from cache, and the cost for these cached inputs is just a small portion of that for uncached inputs. The answers in the tests were scored based on independently prepared reference results; however, since each condition is repeated only a limited number of times, individual percentage point differences cannot be interpreted as precise patterns, and the line graph provided in the article is not a complete public representation of the original data from each run.
From experimental results to a product, boundaries must still be maintained.
According to Asana, the relevant browser navigation changes have been incorporated into the StackAI product; the team is still developing tools that make it more convenient to repeat such experiments. The Command platform, which records requests, trajectories, and results, serves as a repository of evidence here: experimental findings are converted into tickets, undergo code reviews, and are then released, rather than relying on a seemingly successful proxy response to directly determine deployment. For teams that wish to replicate this approach, it is more reliable to first check whether their own request prefixes change frequently, and then compare cache hits, per-request costs, total steps, accuracy rates, and failure rates, rather than simply copying the figure of "20 screenshots." The optimal thresholds also change when cache prices, model context, and task lengths vary.
The case also mentioned an unusual finding: in this task, without any screenshot cropping at all, the cost of a single call for some models was even lower than that of the optimal batching approach. This does not mean that cropping is never necessary; long-term operation will encounter issues such as context limits, drift, and cache invalidation, so there are still limits on steps, tokens, and total costs. The original study did not test all real customer workflows, nor did it prove that the average bill after large-scale deployment would decrease by the same factor. The real insight is that the cost of using these agents is determined not only by the model's pricing but also by how the input history grows, when it changes, and whether rework is required. If a system cannot retain each request and its result, the team may not even be able to determine where the “cost savings” actually come from.












