What a 69-Page AI Report Does Not Show

Frontier firms are reported 8.3x ahead, yet a table on page 35 shows no statistically significant link to revenue.

Author
AuthorAuthor

AI Engineer · 36+ years in IT · Japanese, based in Manila for 13+ years

What a 69-Page AI Report Does Not Show

On 12 August, OpenAI published two reports on how companies use AI. One is called Enterprise Signals; the other is a joint working paper titled How Organizations Use AI: Evidence from ChatGPT. The latter runs to 69 pages and analyses more than 17 million real ChatGPT Enterprise usage records.

The headline was that the gap between companies using AI heavily and those that are not is widening. Yet a small table on page 35 states that no statistically meaningful relationship was found between revenue per employee and AI usage. That finding comes from OpenAI's own data.

For Japanese companies operating in the Philippines, that single line is usable. When headquarters asks how much money AI will make you, you can answer with a fact: even the industry's largest player has not demonstrated it. This article looks at how to read the report as material for your own adoption decision and for explaining that decision upward.

Part 1: Read → Draw the Implications for Your Company

Start with the facts.

Two documents were published. OpenAI's announcement page is dated 12 August, while Fortune records the 69-page working paper as published on 11 August. The paper has five authors: three are OpenAI employees and two are academics. The two academics are David Holtz of Columbia Business School and Prasanna Tambe of the Wharton School at the University of Pennsylvania, and a footnote in the paper states explicitly that both contributed in their capacity as paid contractors for OpenAI.

The figures the report leads with are these. Companies in the top 10 percent of monthly AI usage now generate 8.3 times as many output tokens per active user as typical firms. In January the figure was 2.6 times, so the gap has more than tripled in roughly six months.

The nature of that usage has changed too. As of June, Codex accounted for 64 percent of the combined Codex and ChatGPT output tokens among enterprise customers. Growth in weekly active users since February was 108x in legal, 41x in sales, 41x in recruiting and 26x in marketing, far outpacing the 5x in engineering. The growth is happening in departments that were not using these tools before, rather than in those that already were.

Here are the main figures the report puts front and centre.

MetricFigureBasis of comparison
Output tokens at frontier firms (per active user)8.3xRatio against typical firms
Same figure in January2.6xRatio against typical firms
Share of output tokens from Codex (as of June)64%Codex and ChatGPT combined
Weekly Plugins users (frontier firms)21%Of weekly active users
Weekly Plugins users (typical firms)9%Of weekly active users
Weekly Plugins users (inside OpenAI)95%Of all employees

Growth in weekly Codex users by department since February 2026 is as follows.

DepartmentGrowth
Legal108x
Sales41x
Recruiting41x
Marketing26x
Engineering5x

And then there is the table on page 35. It states that revenue per employee is not meaningfully associated with output tokens per employee or with messages per active user once other controls are included. In other words, no statistically significant correlation could be confirmed.

There is an important caveat attached. The revenue figures used in this analysis are from before employees began using ChatGPT. This is not an analysis that tracked what happened to revenue after adoption. Fortune points out that OpenAI holds data on both usage and revenue, and yet did not extend the study to cover post-adoption changes.

A second chart, on page 29, also matters in practice. AI usage thins out as seniority rises, and the most senior employees recorded the fewest weekly messages per user. Early-career staff use it most. Sarah Friar, OpenAI's CFO, commented on this point in a LinkedIn post to the effect that competitive advantage comes from the people closest to the work.

The chart on page 26 shows total output token growth flattening from around October 2025 through December 2025, then turning sharply upward again from January 2026. The chart ends in March 2026.

Related: Anthropic Reported in Talks Over a $6 Billion Acquisition: How to Handle News That Is Not Settled Yet | Case Study for Japanese Companies in the Philippines explains this in detail.

Part 2: Key Terms for Executives

A few terms are worth pinning down before you read the report, because they are easy to misread.

What "no correlation" actually means In statistics, saying there is no significant correlation is not the same as proving there is no relationship. It reports that, with this data and this method, a relationship could not be demonstrated. Reading it as "it does not work, so do not use it" goes too far in the other direction. The accurate reading is that they wrote down what they did not know, as not known.

Separating what this report can and cannot tell you gives the following.

ClaimHow the report treats it
The depth of AI use is diverging between companiesShown in the data
Usage is spreading beyond technical departmentsShown in the data
Usage thins out as seniority risesShown in the data
Relationship between pre-adoption revenue and AI usageNo statistically meaningful relationship confirmed
What happened to revenue after adoptionOutside the scope of the analysis
Whether more AI usage produces more revenueCannot be determined

What "controlling for other factors" involves It is a statistical procedure: comparing like with like by holding constant other conditions that affect revenue, such as company size and industry. The report notes that before those controls are applied, a pattern is visible in which more profitable companies adopted AI earlier. But that may simply mean profitable companies had the budget to spend on it. Which is cause and which is effect cannot be settled by this analysis.

Output tokens as a yardstick This is the unit used to count the volume of text or code an AI generates. It serves as a proxy for the depth of use. High token volume and real business results are, however, different things. Generating a great deal of long text means nothing if nobody uses it.

What "frontier firms" refers to The report uses this label for companies in the top 10 percent of monthly AI usage. It does not denote a particular industry or company size. By the very definition of a top decile, the membership of that group can change every month.

Related: Lessons from Google's Internal "AI Adoption Gap" Debate: The "Thinking You're Using It" Trap Japanese Firms in the Philippines Fall Into explains this in detail.

Part 3: Applying This to Your Company

Brought down to the day-to-day reality of a Japanese company in the Philippines, this report has three uses.

The first is rebuilding how you explain things to headquarters. If you try to secure budget by arguing that adopting AI will raise revenue per employee, you will not survive the moment someone asks for evidence. This report shows that even the largest player in the industry has not been able to demonstrate it. What you can use instead are figures you can measure inside your own company: the time a process used to take, the number of items sent back for rework, the share of enquiries resolved at first contact. All of these are countable in-house.

When I ran an SEO business in Japan in the 2000s, I spent an hour every day checking search rankings and a full day producing the monthly report. I decided to spend the first two hours of the morning on improvements, reorganised the workflow around that, and cut the daily work down to a third. The ranking-check automation itself, however, lost accuracy every time the search engines changed how they worked, and in the end I went back to manual verification. The cause was that I had not designed it to be repairable when outside rules shifted. What let me explain the benefit at the time was not the automation itself, but the fact that I held a number I could count myself: hours of work.

The second use is building the measurement into the adoption from day one. The weakness of this report is that it only looks at pre-adoption revenue. The same failure happens inside companies. You adopt something, someone asks how it went, and there is no before-figure to compare against. Simply capturing the current numbers for that process in the week you decide to adopt will make the conversation six months later a completely different one.

The third use is deciding who actually uses it. The finding that usage thins out with seniority is, I think, particularly likely to hold in Japanese-affiliated companies. If the people signing off are not using it, they cannot evaluate what the field reports back. Expecting executives to use it daily is not realistic either. What is realistic is that the staff doing the work use it, and that once a month that person presents the actual inputs and outputs and explains them.

Philippine conditions add another layer. A local subsidiary typically has a handful of Japanese expatriate staff and a large majority of local staff. Apply the report's finding that early-career employees use AI most, and the people actually operating it will be your local staff. Making sure it can be run in English, and that operations do not stop when an expatriate is reassigned, are the conditions for it sticking.

And the strong growth in legal, sales, recruiting and marketing points to the departments where a Philippine subsidiary can most easily find an entry point. It means you can start without engineers. First-pass contract checks, matching job descriptions against applications, drafting replies to enquiries — all of these already exist in the daily work of a local subsidiary.

I have supported AI adoption at a Japanese restaurant of about 80 seats near Little Tokyo in Makati. We rebuilt an outdated website in WordPress, targeted small keywords with current SEO and GEO practice, and used AI in social media marketing centred on Facebook. We put an AI chatbot on the site so that simple enquiries are handled automatically, and set the menu up so it can be edited easily from a PC. The consultation took about a week and the build one to two months. What I took from it is that even a site with no engineers can start, provided the entry point is enquiry handling and content updates.

Related: Preparing for Rising AI Prices: AI Cost Management for Japanese Companies in the Philippines explains this in detail.

Part 4: Common Failure Patterns (NG List)

NG 1: Forwarding only the headline to headquarters Send only the figure that the AI usage gap has widened to 8.3x, and the conversation turns into one about how far behind you are. Pass on the other part of the same report as well: that no relationship with revenue could be confirmed. Hand over one half and you will have nothing to say the next time someone asks about return on investment.

NG 2: Using "no correlation" as a reason to hold off This is the misreading in the opposite direction. What was established is only that no relationship was visible between pre-adoption revenue and subsequent usage. What adoption itself produces was outside the scope of the analysis. On its own, it is neither a reason to hold off nor a reason to proceed.

NG 3: Starting without capturing the before-figures This is the most common failure of all. You want to measure the effect, and there is nothing to measure it against. Pick just one of processing volume, elapsed time or rework rate, and record its value on the day you decide to adopt. Numbers reconstructed after the fact always get picked apart in a headquarters meeting.

NG 4: Making token volume or usage counts your success metric Increased usage is not in itself a result. Even the report treats token volume as a proxy for depth of use. Turn it into an internal performance metric and using the tool becomes the objective, which produces a great deal of pointless generation.

NG 5: Signing off without ever touching it The pattern of usage thinning with seniority feeds directly into the quality of decisions. When the field says it is working, management cannot evaluate the substance of that claim. Even setting aside one hour a month to look at real inputs and outputs changes the accuracy of those decisions.

NG 6: Copying sources into internal documents without checking them This report notes in a footnote that two of its five authors are paid contractors for OpenAI. That does not mean the content is wrong, but the premise — that this is a document written by the vendor about its own product — should be shared. Always attach the publication date and the source URL to internal materials.

Practical Tips (3 Tips)

Tip 1: Capture one page of "before-figures" in the week you decide One of processing volume, elapsed time or rework rate is enough to start. Record the pre-adoption value with the date attached. The question this report could not answer is precisely the one created by not tracking those before-figures through to after adoption. It is the cheapest possible way to avoid digging the same hole yourself.

Tip 2: Build the headquarters case on one of your own processes, not on an industry average Figures from external reports make poor grounds for a budget request. With no confirmed correlation, they collapse the moment anyone pushes back. Showing how much time was saved in one of your own processes is stronger material. Even a small number carries more weight when you counted it yourself.

Tip 3: Create a monthly slot where local staff do the explaining The pattern of usage thinning with seniority is, I think, likely to hold in Japanese-affiliated subsidiaries as well. Set aside time once a month for the local staff who actually use the tool to walk through real inputs and outputs. It prevents management from continuing to sign off without seeing the substance. Set it up so it runs in English, and it will survive an expatriate reassignment.

Bonus: How to Make Use of PH AI Works

There are points in working out how to measure the effect of adoption, and how to explain it to headquarters, where an outside perspective helps. PH AI Works offers a 30-minute free consultation for Japanese-affiliated companies operating in the Philippines.

We will work through it with you against the realities of your own operation: choosing the metrics to measure, capturing the pre-adoption figures, structuring automation with tools such as n8n, and preparing the explanatory materials for headquarters in Japan. It is perfectly fine to come while you are still undecided about adopting anything.

References

About the author

Author
Author

Founder / AI Engineer (36+ years in IT)

  • From Tokyo · based in Manila for 13+ years
  • 36+ years in IT (development, SEO, AI)
  • IBM Certified Generative AI Engineer
  • AI chatbots, RAG & AI agent development

A Japanese AI engineer with 36+ years in IT and 13+ years on the ground in the Philippines. I write from hands-on experience to help Japanese companies adopt AI that actually delivers results — chatbots, workflow automation, AI agents, and AI-driven marketing. Feel free to reach out in Japanese or English.

Your Competitors Are Already Using AI!

Is your business keeping up?