Paul Graham Says Token Counts Fail to Measure AI Inference

Industry
โดย Yahoo Finance·USCN·Read original
Summary · why it matters

Y Combinator co-founder Paul Graham argued that token counts are an inadequate way to measure AI inference, saying more capable models solve harder problems with fewer tokens. In a Tuesday post on X, Graham wrote that although you pay for AI by the token, that is not the unit of inference, because you get more problem-solving per token as models improve. He called on the industry to define a standardized unit, suggesting it could involve a chain of increasingly hard problems, each pair of which can be solved by a single model. The argument follows an AlphaSense study last month of 246 tasks that found higher-priced frontier models from OpenAI and Anthropic could deliver better results at lower overall costs than cheaper Chinese models on complex financial analysis tasks. In June, Anthropic launched Claude Fable 5 at twice the price of Claude Opus 4.8, charging $10 per million input tokens and $50 per million output tokens, while Wells Fargo strategist Ohsung Kwon warned that rising AI inference costs could pressure the AI investment trade. In July, Chinese AI models captured a record 58% of tokens processed by U.S. firms on OpenRouter, nearly tripling their share since mid-January, with DeepSeek emerging as the most-used Chinese AI model among U.S. firms.

Impact on stocks 2

Artificial Intelligence · 1 stocks
Financials · 1 stocks

Theme Impact 1

Off-coverage companies 5

AlphaSensePrivate± Mixed
relevance

AnthropicPrivate± Mixed
relevance

DeepSeekPrivate± Mixed
relevance

OpenAIPrivate± Mixed
relevance

OpenRouterPrivate± Mixed
relevance

Related news

2

Google confirms Gemini unintentionally accessed systems at three real companies during security testing

Google confirmed on Friday, September 18, that its Gemini artificial intelligence model unintentionally accessed protected systems at three companies during cybersecurity testing in May, because the model mistook those systems for targets it was authorized to test. Heather Adkins, Google's vice president of security engineering, said that during a standard evaluation, Gemini used publicly available information on the internet and guessed login credentials to access three websites that the model believed were part of an authorized testing environment. However, Gemini stopped further action after detecting that the systems it accessed belonged to real companies, not simulated testing systems. Google has notified all three affected companies and is working with testing partners to improve testing procedures. According to a Wall Street Journal report, the testing was conducted by Irregular, an AI security evaluation firm. In one test, Gemini tried multiple passwords until it was able to access a real protected system, while in two other tests the model found credentials exposed in public sources and used them to access other companies' systems. Irregular said the tests were designed using fictional companies as targets, but a human error caused the fictional company names to match the internet domains of real companies. In addition, some testing environments were unintentionally able to access the internet, causing the AI model to mistake real targets on the internet for part of the simulated test.
InfoQuest·29mRead more →
2

Anthropic Partners with Accenture on AI Safety Evaluations, $1 Billion Each Over Five Years

Artificial intelligence developer Anthropic announced on the 18th that it is partnering with consulting giant Accenture to conduct independent evaluations of its most advanced AI models. Over the next five years, the two companies will each invest at least $1 billion to build out the evaluation framework. Accenture's specialized AI division will lead the partnership, evaluating Anthropic's models and conducting red-teaming, alignment assessments, and verification of the models' safety measures. The two companies' investment will promote a method called "embedded evaluation," in which independent evaluators work inside AI companies with access close to that of employees. Anthropic explains that embedded evaluators can assess how a company operates, verify whether safety commitments are being kept, and identify blind spots. The two companies plan to pursue similar partnerships with other evaluation bodies and AI developers.
ロイター·1hRead more →

Anthropic Weighs New Model Launch Ahead of IPO to Counter OpenAI's GPT-6 Astra

Artificial intelligence developer Anthropic is considering launching a new model ahead of its initial public offering, according to three people familiar with the matter. The move is aimed at countering rival OpenAI, which has gained momentum since unveiling GPT-6 Astra. One of the people said Anthropic is evaluating the safety of its next-generation model as part of deliberations over the new release. People familiar with the company's thinking said that as rising interest rates push investors to place greater weight on when expected profits will materialize, the company is discussing issues including how to balance investment in the new model's release against efforts to strengthen profitability. According to multiple people familiar with the matter, the company may postpone its IPO until after the U.S. midterm elections in November. Anthropic declined to comment.
ロイター·2hRead more →