Wall Street Journal cited views from Goldman Sachs' trading desk, stating that on August AI, the number of calls to the model continued to rise, but the weighted price calculated per million Token fell below $1. This indicates that the growth in demand has not been translated into dollar revenue accordingly, and part of the revenue narrative that has supported the valuation of the AI sector over the past two years is under pressure.
Prices continued to decline in August.

The head of the trading desk at Goldman Sachs, One-Delta, cited in a report Sil the term Data for token issuance. This index fell by 29% in August alone, closing at around $0.97 at the end of the month, which is more than a 50% decline from its high of $2.05 in May. This is the first time that the indicator has fallen below $1 per million Token.
The report also mentioned that on the OpenRouter platform in August, there was a month-over-month increase in usage of approximately 47%, but the expenditure in US dollars only increased by about 7%. In Goldman Sachs' view, what the market is truly concerned about is not the number of calls, but whether these calls will ultimately generate sufficient revenue to cover the costs of data center construction, depreciation, and financing.
Growth in demand has not led to an increase in revenue.
The article argues that a decrease in the price of Token does not necessarily mean a reduction in the usage of AI. It is more likely that model manufacturers reduce prices, users switch to cheaper open-source models, or both scenarios occur simultaneously. The problem is that even when the rate of price decline exceeds the growth in usage, revenue from inference services may still slow down.
Goldman Sachs' trading desk questioned the assumption that "more Token equals more revenue" based on this. If corporate clients can obtain similar model capabilities at lower prices, the unit economics of cloud-based inference will be compressed, and the valuation logic that was originally based on high-growth revenue expectations will also weaken accordingly.
Intensifying Local Reasoning and Competition
The report also attributes the pressure to two changes. First, there has been an improvement in local inference capabilities; some high-performance laptops, workstations, and devices equipped with NPU are now capable of running more powerful models. As costs continue to rise in API, companies may opt to make one-time purchases of hardware. Second, competition among models has intensified, with cheaper alternatives emerging continuously, which undermines the premium space for leading closed-source models.
In this context, the model that charges according to Token faces greater challenges. The article argues that for this model to be sustainable in the long term, it would require at least models that are difficult to run locally, a lack of sufficient open-source alternatives, and workloads that are sufficiently volatile, but all of these conditions are weakening.
Capital expenditure is first tested by the credit market.
Goldman Sachs' trading desk also turned its focus to the capital expenditures of hyperscalers in the cloud industry. The report states that the five rated cloud providers are expected to spend approximately $737 billion in capital expenditures in 2026, which accounts for about 38% of their revenue. If revenue growth slows down while depreciation and interest costs continue to rise, return on investment will be under pressure.
The article mentions that the credit market has already reflected this concern ahead of the stock market. Statistics show that among the 91 bonds issued by relevant cloud vendors in 2026, 78 had fallen below their issue prices by the end of August. Based on this, Goldman Sachs' trading desk believes that the bond market is re-evaluating the return prospects of AI's heavy asset expansion.
As for the variables that may change short-term expectations, the report considers the next-generation model Astra under OpenAI to be one of the few catalysts. However, the article also points out that if high-value workloads remain in limited testing environments and do not enter large-scale public billing systems, then the new model may not immediately improve revenue performance.











