Olas Unveils Small AI Model That Matches GPT-4.1 in Forecasting Test

โดย Benzinga·US·Read original
Summary · why it matters

Olas unveiled a new AI model Tuesday that was trained on more than 200,000 prediction market forecasts and matched OpenAI's GPT-4.1 in a test of its ability to predict real-world events. Olas-Predict-R1-14B achieved 75.8% accuracy across 2,628 previously unseen markets, compared with 75.4% for GPT-4.1 and 71.4% for the underlying DeepSeek model, according to benchmark materials shared with Benzinga. Olas, which develops autonomous AI agents that research and trade on prediction markets, trained the model using 214,529 forecasts from 5,116 resolved markets, and fine-tuning improved the model's Brier score by roughly 20%. David Minarsch, CEO of Valory and founding member of Olas, told Benzinga the results suggest cheaper, specialized AI models are putting increasing economic pressure on the industry's most advanced general-purpose systems, and said frontier models may need to find new markets to maintain their growth rates. Olas says its model can run on a single GPU and its weights are publicly available, letting developers run it themselves instead of paying a closed AI provider for each forecast. In a separate experiment requested by Benzinga, Olas-Predict gave Nvidia a 70% chance of ending 2026 as the world's largest company by market capitalization, while Polymarket traders currently put Nvidia at 78%.

Impact on assets 3

Artificial Intelligence · 2 stocks
Others · 1 stocks

Theme Impact 3

Off-coverage companies 3

DeepSeekPrivate± Mixed
relevance

OpenAIPrivate± Mixed
relevance

PolymarketPrivate± Mixed
relevance

Related news

6impact 4

OpenAI Seeks $1.2 Trillion Valuation in New Funding Round

OpenAI, the AI partner of Microsoft, is reportedly seeking a valuation above US$1.2 trillion in a new funding round. The AI group is forecast to burn roughly US$278 billion in cash as it scales infrastructure and product development with Microsoft support. The prospective deal would reinforce OpenAI's role inside Microsoft's broader AI ecosystem, spanning Azure infrastructure and Copilot branded services. The funding push backs the view that rich AI workloads can keep Azure traffic and Copilot adoption growing, while also sharpening questions about capital intensity and dependency on a few large AI partners.
Simply Wall St·52mRead more →
impact 4

China Opens Probe Into DeepSeek and Moonshot AI Over Suspected Leak of Classified Data

China's Cyberspace Administration has launched an investigation into Chinese startups DeepSeek and Moonshot AI over suspicions that state secrets were leaked to U.S. AI companies. The U.S. tech outlet The Information reported on the 22nd that Chinese authorities began the probe amid concerns over "distillation," in which the outputs of high-performance American AI models are used to train Chinese AI systems. According to a report published by U.S. AI startup Anthropic, DeepSeek and Moonshot sent users' questions to Anthropic's cutting-edge Claude model without authorization and used its answers to train their own models. The data involved also included internal information from the Chinese People's Liberation Army, public security authorities, and state-owned enterprises. The Trump administration and U.S. national security and investigative authorities have sharply condemned distillation. Ahead of the U.S.-China summit on the 24th, the issue of data security has come into sharp relief.
Jiji Press·1hRead more →
3

Anthropic Unveils Claude Opus 5.5 as OpenAI Releases Terra and Luna Models

Anthropic unveiled its new Claude Opus 5.5 on Tuesday, a model the developer says offers performance approaching its frontier Fable 5.1 at a lower cost for enterprise customers. OpenAI also just released its new GPT-6 Sol and Luna models, which are its lower-end models outside of Solar and Astra that it previously had. The twin launches underscore a shift across the AI industry toward cheaper offerings as companies rein in spending on tokens and shop around for models that fit their needs and budgets. Anthropic included notes on security in its release announcement, while OpenAI has said it is working with third-party external partners to audit its AI models. The push comes amid debate over AI safety and a lawsuit alleging a group of large companies colluded to slow AI development and degrade performance for paying consumers.
Yahoo Finance·1hRead more →