Inference-Optimized Silicon

The AI chip the world knows (NVIDIA's GPU) is built to be great at "training" a model. But training happens only a few times. Running the model to actually answer real questions (inference) happens every second, millions and billions of times. Once the volume of running dwarfs training, the cost per answer (cost per token) and the speed (tokens per second) become the new battleground — and that opens the door to a new breed of chip that does "only the running," faster and cheaper than a general GPU. This lesson introduces Groq, Cerebras, d-Matrix, and the reason most of the real players are still private companies.

Theme index · base 100 · USD total return
News & notes moving Inference-Optimized Silicon
Inference-Optimized Silicon

Xeal Launches Laitent, World's First Edge Inference Compute Network Using Idle EV Charging Capacity

Xeal launched Laitent, which it calls the world's first edge inference compute network using idle EV charging capacity, tapping more than 200MW of permitted, installed electrical infrastructure across 1,600+ properties. Xeal, a member of NVIDIA Inception, plans to deploy over 100,000 NVIDIA GPUs alongside EV charging infrastructure, and has secured partnerships with Rafay Systems for AI infrastructure orchestration, Spectrum Business for dedicated enterprise-grade fiber, dozens of real estate and property managers, and a Tier 1 inference provider for up to 5MW of compute. The first Laitent Pod will be brought online with partner JVM Realty by the end of 2026. Each Laitent Pod is about the size of one parking space, contains up to 48 NVIDIA Hopper or Blackwell Ultra GPUs, requires no water hookup, and runs quiet at less than 65 decibels, offering sub-20ms latency in metro areas. Xeal said EV charging sites typically operate at less than 10% of permitted capacity, and it taps the remaining 90% for compute, with property owners able to add as much as $1m in property value for little-to-no upfront investment. Looking beyond the initial 200MW of installed charging capacity, Xeal plans to unlock over 1GW of existing headroom across real estate and EV charging deployments.
Business Wire·16hRead more →
Inference-Optimized Silicon5impact 4

Nvidia CEO Jensen Huang Sees Chip Sales Doubling in 2027

Nvidia CEO Jensen Huang said the company expects chip sales next year to be about twice this year's level, sending shares up more than 2% Thursday. Speaking at an event in Scotland, Huang pointed to continued demand as businesses expand their use of artificial intelligence. Nvidia also released preliminary MLPerf results on Sept. 16 showing its next-generation Vera Rubin NVL72 platform delivering up to 3.7 times the inference throughput of the previous GB300 system on the Qwen3-VL test. Vera Rubin has entered full production, with shipments expected to begin this fall. The broader market added to the lift, as U.S. stocks rebounded Thursday with oil prices falling more than 2% and the 10-year Treasury yield easing, helping technology shares recover from recent pressure.
GuruFocus·1dRead more →
Inference-Optimized Silicon2impact 4

Qualcomm Gives Amazon Warrants Tied to $60 Billion AI Chip Deal

Qualcomm has given Amazon an equity-linked incentive to deepen their AI-infrastructure partnership, with Amazon able to eventually purchase as much as $60 billion of Qualcomm data-center products while purchase-linked warrants could hand Amazon roughly $4 billion of Qualcomm shares at $161.26 each. Qualcomm shares gained approximately 2.1% to $188.64 Thursday, a level 7.37% above the stock's $175.69 GF Value. The warrant value represents roughly 6.7% of the maximum purchasing framework, directly tying Amazon's buying activity to Qualcomm's equity story. The two companies are also working together on custom AI inference silicon and optical connectivity capable of reaching 1.6 terabits per second. No minimum purchase commitment, delivery timetable or margin profile has been disclosed.
GuruFocus·1dRead more →
Inference-Optimized Siliconimpact 4

Meta Could Save $8.5 Billion in 2027 on Custom MTIA Chips, BofA Estimates

Bank of America estimates Meta Platforms could save roughly $8.5 billion in 2027 by running AI workloads on its own custom silicon instead of buying third-party chips, an outside analyst estimate rather than company guidance. The figure rests on a specific roadmap: Meta plans to deploy its third-generation MTIA 450 chip, code-named Arke, in the first half of 2027, followed by the higher-performance MTIA 500, or Astrid, later that year, both co-developed with Broadcom and aimed at AI inference workloads. BofA models Meta deploying 5 to 6 gigawatts of owned capacity in 2027 at a total cost of roughly $200 billion, assumes chips make up 60% of that spend, and pegs Meta's custom silicon as about 40% cheaper than third-party equivalents. Broadcom CEO Hock Tan said custom chips optimized for a customer's own workloads outperform any GPU and can do so at half the cost, and confirmed Broadcom will deliver three generations of MTIA accelerators to Meta between now and the end of 2027. Meta's FY2026 capex guidance sits at $130 billion to $145 billion, narrowed from $125 billion to $145 billion, with total expense guidance raised to $165 billion to $169 billion, while Q2 2026 revenue reached $60.80 billion, up 27.96% year over year, on advertising revenue of $59.36 billion.
Yahoo Finance·1dRead more →
Inference-Optimized Silicon

Intel Posts MLPerf Inference Gains Across Xeon 6 and Arc Pro GPUs

Intel reported performance gains in its latest MLPerf Inference v6.1 results across Intel Xeon 6 processors and Intel Arc Pro B-series GPUs. Intel Xeon 6980P processors delivered 2.4x higher Llama 3.1 8B Server throughput and 56% higher Offline throughput than MLPerf v6.0 on the same hardware setup, while customer and partner results rose from 29 to 39, with Oracle, Red Hat, Quanta Cloud Technology and Supermicro making their first submissions. Intel expanded Xeon 6 participation from two to five processor models, lifting CPU inference results from 24 to 35, and its Arc Pro B70 GPUs supported workloads including Llama 2 70B, gpt-oss-120B and Whisper, with a four-GPU system offering 128GB of Video Random Access Memory and Server performance up 36% and Offline performance up 27% versus MLPerf v6.0. Intel also co-developed results for the new end-to-end retrieval-augmented generation benchmark, using a system that paired a Xeon 6787P processor with four Arc Pro B70 GPUs to split the workload between CPU and GPUs. Intel faces competition from Qualcomm, which is expanding into AI data-center infrastructure through its Dragonfly platform, and from AMD, which is strengthening its AI infrastructure with the Helios platform.
Zacks Investment Research·1dRead more →
Inference-Optimized Silicon

Wall Street Analysts See Bottom Forming in Hammered 2026 IPO Stocks Cerebras and Innio

Wall Street analysts are flagging a potential bottom in two 2026 IPO stocks that have fallen sharply since going public. Cerebras Systems, which started trading on May 14 at $350 and closed its first day at $311, has dropped 41% since then, though Morgan Stanley analyst Joseph Moore rates it Overweight with a $279 price target, implying 52% upside, and the Street's Strong Buy consensus carries a $296 average target versus a current $184.03. The company's $20 billion OpenAI deployment deal, running in stages through 2028, makes up the bulk of its $25.4 billion in remaining performance obligations, though that customer concentration worries some investors. Innio Holding, which went public on June 4 at $27 per share and closed its first session at $33.30, has fallen 46.5% since then, but RBC analyst Chris Dendrinos rates it Outperform with a $35 target, implying 97% upside, and the Strong Buy consensus average target of $40.70 implies 129% upside from the current $17.79. Innio's first public earnings report showed equipment order intake up 316% year-over-year to $2.3 billion and backlog up 279% to $6.6 billion, with total revenue of $937.7 million for 2Q26, a 42% year-over-year gain that beat estimates by $54.37 million.
TipRanks·2dRead more →
Inference-Optimized Silicon4impact 4

Meta to Deploy Custom AI Chips in Data Centers by 2027

Meta Platforms is moving deeper into custom AI silicon, with a new generation of internally designed chips set to enter its data centers in the first half of 2027. The company is currently testing its third-generation MTIA 450 processor, code-named Arke, while its successor, MTIA 500, or Astrid, is expected to complete design work in about a month and reach data centers by the end of 2027. Meta is working with Broadcom on chip design and Taiwan Semiconductor Manufacturing Co. on production, and has committed to deploying more than a gigawatt of the chips over a 12-month period. Twelve Arke chips delivered by TSMC on Sept. 1 performed within 2% to 3% of Meta's simulations and have already run Meta models alongside models from DeepSeek and Alibaba. Meta also canceled its planned Olympus processor, which was intended to handle both AI training and inference, partly because of cost concerns, and is instead prioritizing inference, the day-to-day running of AI models.
GuruFocus·3dRead more →
Inference-Optimized Silicon

Dell Shares Rise 5.8% as Axelera AI's Europa Chip Enters Its Servers

Dell Technologies shares rose about 5.8% to $565.15 on Tuesday after European startup Axelera AI unveiled its Europa inference processor, which Reuters reported is designed for enterprise inference workloads and is expected to appear in future systems from Dell and Supermicro. Axelera says more than 600 customers already use its technology, with signed agreements worth tens of millions of dollars and a broader potential sales pipeline that could reach $1.5 billion, an opportunity estimate rather than booked revenue. Dell's larger AI engine remains its server business, which exited the latest quarter with roughly $95 billion of AI-server backlog after shipping $16.4 billion of AI systems. Even if Axelera captured the full $1.5 billion opportunity, that would equal only about 1.6% of Dell's existing AI backlog. The strategic takeaway is that Europa gives Dell another inference supplier and potentially more power-efficient hardware choices, though Nvidia-based systems still carry far more weight in Dell's near-term AI economics.
GuruFocus·3dRead more →
Inference-Optimized Silicon

Cirrascale Launches Production Release of Enterprise AI Inference Platform

Cirrascale Cloud Services announced the production release of the Cirrascale Inference Platform, a complete software stack for enterprise-grade AI inference, at the AI Infra Summit in Santa Clara. The platform lets enterprises run open-source models, their own private models, and closed model ecosystems from a single serverless platform, including Google Gemini delivered on premises through Google Distributed Cloud and operated by Cirrascale. Its model and hardware selection layer automatically routes each request to the right model and runs it on the best available accelerator across NVIDIA, AMD, Tenstorrent, and Qualcomm with no code changes required to switch, and teams can fine-tune models on their own private data without that data leaving their environment. The platform includes a turnkey private chat experience connected to a company knowledge base, built-in controls to manage AI spend across teams, and governance guardrails for agentic workloads, aligned with HIPAA, SOC 2, and FedRAMP requirements where required. The Cirrascale Inference Platform is available now across Cirrascale's U.S. and international regions.
Business Wire·3dRead more →
Inference-Optimized Silicon6impact 4

Broadcom CEO Hock Tan Reaffirms AI Outlook as Chip Stocks Slide

Broadcom CEO Hock Tan said he sees little reason to change the company's longer-term AI outlook, pushing back on fears of a slowdown in frontier artificial intelligence development that sent chip stocks sharply lower on Monday. Speaking in a Monday CNBC interview, Tan said Broadcom still expects demand for computing infrastructure used for both AI model development and inference to remain durable, remarks that reinforced the company's existing forecasts for fiscal 2027 and 2028. The selloff followed comments from Anthropic CEO Dario Amodei calling for a more measured approach to advances in AI models, with OpenAI CEO Sam Altman and Elon Musk also raising concerns about the pace and risks of AI development. Tan expects Anthropic to become Broadcom's largest custom-chip customer in 2027 and to retain that position in 2028, overtaking Alphabet as the company's biggest custom-chip customer. He also pointed to inference, the running of trained AI models for everyday use, as another potential source of sustained demand.
GuruFocus·3dRead more →
Inference-Optimized Silicon

AnalogAI Licenses Microchip's SST memBrain SAGE IP for Edge AI Processors

AnalogAI has selected the memBrain Synaptic Analog Generative Engine neuromorphic hardware intellectual property from Microchip Technology's Silicon Storage Technology subsidiary for its first real-world edge AI processors. AnalogAI is using the SST memBrain SAGE IP to deliver analog compute-in-memory performance at or below one watt for ultra-low-power edge applications, targeting environment-adapting humanoid robots, drones and vehicles. Mark Reiten, senior vice president of Microchip's Intelligent Compute business unit, said the IP serves as the core inference engine for AnalogAI's first products, delivering the compute performance and power efficiency the company requires. AnalogAI chief executive officer Jaejun Lee said the company chose the silicon-proven memBrain SAGE IP after an industry-wide search because it accelerates development while meeting ultra-low-power and high-performance targets. The memBrain SAGE IP has been developed and deployed in 40 nm and 28 nm foundry processes using production-ready SuperFlash memory, with a roadmap that includes 22 nm development.
GlobeNewswire·3dRead more →
Inference-Optimized Silicon4

Euclyd raises $231 million Series A with Samsung backing

Dutch semiconductor startup Euclyd has raised $231 million in a Series A round backed by Samsung. The 200-million-euro round was co-led by Samsung alongside Somerset Capital Partners, the Scaleup Europe Fund managed by EQT, and Innovation Industries, according to CNBC. Founded in 2024 by Bernardo Kastrup and Atul Sinha at High Tech Campus Eindhoven in the Netherlands, Euclyd is building chip systems designed to reduce the energy consumption and cost of running AI models, targeting AI inference rather than model training. Chief executive officer Bernardo Kastrup told CNBC that Samsung brings more than capital, citing its memory manufacturing, engineering depth, systems knowledge, supply chain and network. Dede Goldschmidt, senior vice president of Samsung Electronics and head of the Samsung Semiconductor Innovation Center, said Euclyd's approach addresses real constraints in AI data centers. The company plans to use the proceeds to expand its engineering team, advance its silicon and systems development, and build out ecosystem partnerships, and Kastrup said it expects to launch its physical chip systems in 2028 and serve thousands of enterprise customers by 2030, though Euclyd has not yet demonstrated its chip systems in large-scale commercial environments. The funding arrives as investment in AI chip alternatives to Nvidia has grown, with Nvidia's share of the inference chip segment estimated between 60% and 75% in 2025 facing growing pressure from custom silicon developed by hyperscalers and startups alike.
Inference-Optimized Silicon

EUCLYD Raises Over €200 Million Series A to Break AI Efficiency Wall

EUCLYD has signed a Series A financing round of more than €200 million, the Eindhoven-based semiconductor systems company announced. The round is co-led by Samsung, Somerset Capital Partners, Scaleup Europe Fund, which is managed by EQT, and Innovation Industries, with participation from EIFO, the Export and Investment Fund of Denmark, imec.xpand, Brabant Development Agency (BOM), and Quadri. Peter Wennink, former President and CEO of ASML, joins EUCLYD as Chairman of the Board. The financing will expand EUCLYD's engineering organization, accelerate its silicon and systems roadmap, strengthen ecosystem partnerships, and prepare the company for commercial deployment across enterprise, sovereign, and hyperscale AI markets. At the center of EUCLYD's roadmap are craftwerk, which it describes as the world's first agentic AI silicon, and craftwerk station CWS, which it calls the world's lowest-power exascale AI factory.
PR Newswire·4dRead more →
Inference-Optimized Silicon

Intel Could Gain as AI Spending Shifts From Training to Inference

A September 14 selloff treated slower frontier AI development as bad news for chips, but Reuters Breakingviews argued that a moderation in giant model-training runs could redirect part of this year's roughly $1 trillion AI investment toward inference, where existing models answer queries and run agents. That shift would create a different contest for Intel and Nvidia, because inference does not erase GPUs but broadens the bill of materials. Intel's opportunity is the server CPU, since production AI applications need orchestration, memory, networking and general-purpose compute around accelerators, and Intel reported $6.3 billion of Data Center and AI revenue in its latest quarter, up 59% year over year. Execution remains the risk: Intel is spending heavily to rebuild manufacturing leadership and still generated negative adjusted free cash flow in its latest quarter, and inference growth helps only if Intel keeps enough share and margin against AMD, Arm-based CPUs and specialized accelerators. Nvidia remains the strongest counterargument to a simple rotation thesis, since its GPUs run both training and inference and CUDA plus newer inference-focused systems give it a path to monetize deployed AI even if frontier training slows, with the downside relative rather than existential. Insider Monkey's database counted 138 hedge funds holding Intel in Q2 2026, up from 112 in Q1, with AQR Capital Management owning 10,742,567 shares after trimming its position 7%, while Nvidia rose to 285 funds from 275 and Fisher Asset Management increased its stake 3% to 90,935,947 shares, filings that predate the September 14 safety-driven market reaction. Intel had 152,248,752 shares sold short on August 31, equal to 3.02% of float, with 1.69 days to cover.
Insider Monkey·4dRead more →
Inference-Optimized Siliconimpact 4

Cerebras Expands AI Inference Push Against Nvidia and Alphabet

Cerebras Systems is expanding its technology roadmap and customer base to strengthen its position in the AI inference market, targeting coding, agentic AI, security and enterprise workloads as it competes with NVIDIA and Alphabet. Core cloud and other services revenues surged 287% year over year to $127.7 million in the second quarter of 2026. The company expects to double system speed annually over the next several years and increase throughput by more than 20 times over the next 18 months, while its disaggregated inference architecture with AMD combines Cerebras systems with AMD Helios racks to raise throughput by up to five times. Its collaboration with Amazon Web Services is expected to bring disaggregated inference to Amazon Bedrock in the first quarter of 2027, and Cerebras supports OpenAI's GPT-5.6 Sol at a speed that is 10 times faster, with first revenues from other hyperscalers expected around mid-2027. Cerebras signed six deals worth more than $30 million each in the second quarter and added customers such as Cognition, Lovable and Figma in AI coding, while Block, AlphaSense and GSK use its fast inference for agentic workflows and CrowdStrike represents a move into AI-powered cybersecurity. The company has more than 600 MW of data-center capacity live or contracted through the end of 2027, and manufacturing capacity is expected to rise more than tenfold in 2026, supported by secured TSMC wafer supply.
Zacks Investment Research·4dRead more →
Inference-Optimized Silicon3

Piper Sandler Sees Intel Revenue Growth in High Teens Through 2030

Piper Sandler initiated coverage of Intel with a neutral rating and a $110 price target, even as the firm projects the chipmaker can sustain annual revenue growth in the high teens through 2030. Intel's top line was flat last year at $52.9 billion, and consensus estimates now project a 19% increase in 2026 to $63 billion. The expected acceleration is driven by rising server CPU demand for AI data centers running inference and agentic workloads, a shift that Advanced Micro Devices estimates is moving the server CPU-to-GPU ratio toward 1:1 from 1:4 or 1:8. Intel is reportedly sitting on a server CPU order backlog of over six months, and server CPU prices have already risen 10% to 35%, with DigiTimes reporting the company plans a further 10% price increase next month. Bank of America expects the server CPU total addressable market to grow almost 5x between 2025 and 2030, reaching over $170 billion by the end of the decade.
The Motley Fool·5dRead more →
Inference-Optimized Siliconimpact 4

DeepSeek Cuts KV-Cache Memory Needs, Pressuring Micron and Sandisk

DeepSeek's September 10 release says its V4.1-Flash model needs one-quarter as much HBM and one-eighth as much SSD capacity for its KV cache as the previous generation, a 75% cut in HBM and an 87.5% cut in SSD for that component. The comparison covers only the KV cache, which stores attention state during inference, and is measured against DeepSeek's preceding architecture; it does not cover memory used for model weights or training. The model has 552 billion parameters but activates 8 billion for input and 16 billion for output. Micron's Cloud Memory revenue reached $13.77 billion in its latest quarter at an 83% gross margin, while Core Data Center revenue was $11.52 billion at an 87% margin, and the company shipped more than $1 billion of HBM4. Sandisk's quarterly data-center revenue rose 103% sequentially to $2.98 billion, with pricing generating roughly two-thirds of its companywide sequential revenue increase. As of August 31, 5,996,105 Sandisk shares were sold short, equal to 4.10% of float and 0.46 days of average volume.
Insider Monkey·6dRead more →
Inference-Optimized Silicon

Nvidia Opens NVLink Fusion to d-Matrix Raptor Chips, Astera Labs Aids Connectivity

Nvidia is letting rival AI accelerators into its racks, as Reuters reported September 10 that inference-chip startup d-Matrix will use Nvidia's NVLink Fusion technology to connect its next-generation Raptor processors directly into Nvidia data-center racks, with the systems expected to become available in 2027. Astera Labs is working with d-Matrix on custom connectivity solutions to move data rapidly across the system, creating a three-layer architecture in which d-Matrix supplies the accelerator, Nvidia supplies the rack-scale interconnect technology and architecture, and Astera helps solve the data-movement problem between components. For Nvidia, the bull case is that NVLink can become more valuable even when Nvidia does not sell every accelerator, shifting it from dominating GPUs to controlling an important system standard, while the bear case is that NVLink Fusion deliberately makes alternative accelerators easier to deploy and could cost Nvidia some GPU unit share. For Astera Labs, heterogeneous systems may be an even cleaner opportunity, since more accelerators, memory pools, CPUs and switches create more connectivity bottlenecks for its retimers, fabric switches and other connectivity products, though its valuation already assumes substantial AI infrastructure growth. Hedge-fund positioning turned more bullish on Astera Labs in the second quarter, with 73 hedge funds tracked in the stock versus 53 in the first quarter, while Nvidia rose to 285 hedge funds from 275, and short interest stood at roughly 6.6% of Astera's float on August 14 versus about 1.2% for Nvidia, though ALAB short interest fell from the previous reporting period.
Insider Monkey·7dRead more →
Inference-Optimized Silicon

Nvidia and CrowdStrike Unveil SafeMind Agentic Cybersecurity System

Nvidia and CrowdStrike unveiled SafeMind, an agentic cybersecurity system built on Nvidia Nemotron models and technology from CrowdStrike's Cyber Superintelligence Lab, at CrowdStrike's Fal.Con 2026 conference. Nvidia described the architecture as a continuous loop in which offensive and defensive AI systems challenge and improve one another. Nvidia CEO Jensen Huang reinforced the thesis on September 10 at Goldman Sachs' Communacopia + Technology Conference, arguing cybersecurity is a natural successor to AI coding because defensive systems can operate continuously rather than waiting for a human prompt. For CrowdStrike, AI is simultaneously creating a threat and a product opportunity, since attackers can automate vulnerability discovery, phishing, malware development and lateral movement, while CrowdStrike's endpoint and identity telemetry gives its AI agents proprietary context. For Nvidia, cybersecurity creates a potentially attractive inference workload, as security agents may operate continuously across millions of endpoints, generating recurring demand for inference rather than one-time model training. Hedge-fund ownership of CrowdStrike rose to 89 funds in Q2 from 79 in Q1, while Nvidia rose to 285 from 275, and short interest stood at about 2.4% of CrowdStrike's float and 1.2% of Nvidia's float as of August 14.
Insider Monkey·7dRead more →
Inference-Optimized Silicon6impact 4

Qualcomm Secures AWS Data Center Deal With $4 Billion Warrant and $60 Billion Purchase Potential

Qualcomm has secured a multi-generational silicon partnership with AWS, one of the world's biggest AI infrastructure spenders, sending its shares up over 3%. Under the agreement, Qualcomm granted Amazon warrants covering 25 million shares at $161.26 per share, giving Amazon the right to potentially invest roughly $4 billion if the warrants are fully exercised, with vesting tied to commercial arrangements, binding purchase orders and Amazon's purchases of up to $60 billion worth of Qualcomm server chips. The $60 billion figure does not represent committed revenue, though 3.75 million shares vested at issuance based on Amazon's initial purchase commitments, making the total a potential ceiling dependent on successful execution. The deal adds momentum to Qualcomm's data center push following the June unveiling of its Dragonfly C1000 CPU, for which Meta is a named 2028 customer, and supports the company's target of reaching $15 billion in data center sales by fiscal 2029. Qualcomm's strategy focuses on AI inference and CPUs rather than the GPU training market dominated by Nvidia, and in June the company raised its fiscal 2029 non-handset revenue target from $22 billion to $40 billion, a target that remains unvalidated by shipped products and revenue-generating volume at scale.
Insider Monkey·7dRead more →
Inference-Optimized Siliconimpact 4

Cerebras Raises 2026 Revenue Guidance After Q2 Beat, Shares Down 17.2%

Cerebras reported second-quarter 2026 results that beat estimates and raised its full-year core revenue guidance to $880-$890 million, even as its shares have fallen about 17.2% in the month since the report. The company posted a quarterly loss of 4 cents per share, an 80.95% earnings surprise, while core revenues of $209.87 million rose 103% year over year and topped consensus by 8.09%. Core cloud and other services revenues surged 287% year over year to $127.7 million, and remaining performance obligations reached $25.4 billion. For the third quarter of 2026, Cerebras expects core revenues of $214-$216 million, core gross margin of 38% to 40%, and core operating margin between negative 25% and negative 23%. Management said it has secured more than 600 megawatts of data center capacity live or under contract for delivery by the end of 2027, and it expects core revenues to more than triple in 2027.
Zacks Investment Research·7dRead more →
Inference-Optimized Silicon8impact 5

Qualcomm Strikes Amazon Deal Worth Up to $60 Billion for AI Data-Center Chips

Qualcomm has struck a long-term partnership with Amazon under which Amazon could purchase as much as $60 billion of Qualcomm's AI data-center chips and related products. As part of the agreement, Qualcomm granted Amazon warrants worth roughly $4 billion, allowing Amazon to buy Qualcomm shares at $161.26 per share, with the warrants vesting as Amazon purchases Qualcomm products. The two companies will develop custom AI chips focused on AI inference, along with high-speed optical connectivity for data centers, and Qualcomm is targeting $15 billion in annual data-center chip revenue by 2029. Amazon joins Microsoft and Meta as major cloud customers working with Qualcomm on custom AI chips, though the $60 billion figure is a potential purchase commitment rather than guaranteed revenue, and Qualcomm still must commercialize chips that can compete with Nvidia. Amazon's annualized custom-chip revenue run-rate exceeded $25 billion at the end of the June quarter, and Reuters reported separately that Amazon was preparing its first sterling bond sale as hyperscalers seek funding for AI infrastructure costs.
Insider Monkey·7dRead more →
Inference-Optimized Siliconimpact 5

Google, NVIDIA and SpaceX Deals Reshape Global Compute Race

The compute landlord thesis went global this week as Google, NVIDIA and SpaceX each moved to lock down power, distribution and capacity. Google committed €13B to Finland, securing a 22-year power purchase agreement with the Fortum Loviisa nuclear plant. NVIDIA reportedly agreed to acquire Hugging Face for $12.9B, taking control of the main conduit for open-weight models such as Qwen and DeepSeek, which account for 61% of tokens consumed on OpenRouter. At SpaceX, an undisclosed tenant signed a $13.3B annual commitment, lifting total ARR for the hosting unit to roughly $41B across four pillars — Anthropic, Google, Reflection AI and the mystery customer — with 90-day termination clauses starting in 2027. In China, prices for Huawei and Cambricon AI chips are surging 20% to 50% as export controls push manufacturers into grey-market high-bandwidth memory; the Huawei Ascend 950DT now carries an indicated price above 250,000 yuan, roughly $37,000 per accelerator, while Cambricon's forthcoming 690 chip has been repriced 20% to 30% higher than quotes from two months earlier. DeepSeek V4.1 Flash cuts inference costs by 80% through its Causal Encoder-Decoder architecture, compressing cache-hit costs to $0.003 per token, and Positron AI raised an $875M Series C at a $5B valuation for its Asimov chip, which swaps scarce high-bandwidth memory for commodity LPDDR5X and claims 90% bandwidth utilization against NVIDIA's typical 30%.
Yahoo Finance·7dRead more →
Inference-Optimized Siliconimpact 4

Positron AI Raises $875 Million to Challenge NVIDIA's HBM Inference Lock

Positron AI has raised $875 million in a combined Series C and Series C-1 round, valuing the Reno-based startup at $5 billion post-money. The financing, co-led by NEA, Andra Capital, Atreides Management, Valor Equity Partners, and SemiAnalysis Capital, represents a roughly five-fold markup from the company's February 2026 Series B valuation of just over $1 billion. The bet is that AI inference economics can be redesigned by replacing supply-constrained high-bandwidth memory with commodity LPDDR5X, the approach behind Positron's Asimov inference ASIC built on TSMC's N3P process. A single Asimov chip carries 288 GB of on-package LPDDR5X, expandable to 2.3 TB via CXL expansion, and Positron claims greater than 90% bandwidth utilization on transformer inference workloads versus less than 30% typically realized on NVIDIA GPUs. Asimov is scheduled for tapeout by the end of 2026, with production slated for the second half of 2027, meaning the chip will not be commercially available for at least twelve months.
Yahoo Finance·8dRead more →
Inference-Optimized Silicon

Nvidia Falls 2.5% as $2 Billion Startup d-Matrix Joins NVLink Fusion

Nvidia shares slid 2.5% to $217.97 Thursday morning even as $2 billion chip startup d-Matrix joined its NVLink Fusion ecosystem. D-Matrix plans to connect its Raptor inference processors directly to Nvidia's rack-scale infrastructure, widening the reach of Nvidia's technology beyond its own chips. The startup aims to complete Raptor's final design in 2026, with compatible racks arriving in 2027, targeting speed-sensitive applications such as coding assistants, chatbots and voice agents. Investors received no details on financial terms, purchase commitments or expected shipment volumes, leaving the immediate revenue impact unclear. NVLink Fusion lets custom processors and accelerators operate alongside Nvidia's GPUs, networking, memory and rack technology, a strategy that matters as data centers generate roughly 92.5% of the company's latest $96.2 billion quarterly revenue.
GuruFocus·8dRead more →
Inference-Optimized Silicon

d-Matrix Adopts NVIDIA NVLink Fusion for Rackscale AI Inference

d-Matrix announced a multi-year collaboration with NVIDIA that will integrate its next-generation inference XPUs, starting with d-Matrix Raptor, directly into NVIDIA's MGX rackscale architecture via NVLink Fusion. The rack-level system is aimed at AI labs, hyperscalers, and neoclouds seeking ultra-low latency premium-level token services, and pairs d-Matrix XPUs with NVIDIA Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs, and Spectrum-X Ethernet networking. d-Matrix is also partnering with Astera Labs to deliver custom connectivity solutions for high-throughput data flow, and the MGX-based rack will use modular cable-free trays built on NVIDIA's proven supply chain. Raptor, a follow-on to the Corsair XPU platform, uses a first-of-its-kind 3D DRAM stacking approach that combines a DRAM memory chip and an SRAM compute chip into a single two-story package, is backed by more than 100 patents, and is expected to tape out before the end of the year. Initial availability of d-Matrix Raptor XPUs integrated into the NVIDIA MGX rack is expected in Q4 2027.
PR Newswire·8dRead more →
Inference-Optimized Silicon7

Piper Sandler Initiates Five AI Chip Stocks at Overweight

Piper Sandler launched coverage of the artificial intelligence chip sector on Thursday, initiating five semiconductor stocks at Overweight as prime beneficiaries of surging AI compute demand. Analyst David O'Connor started Nvidia with a $300 price target, calling it the "outright leader in AI compute with 80% market share" and describing the stock as "among the cheapest in AI universe" at about 14 times fiscal 2028 estimates. Broadcom was initiated at $460 as the leader in custom ASIC chips with a 75% share, where O'Connor wrote that "demand is 2x supply" and noted line of sight to 12 gigawatts of demand in fiscal 2027. Advanced Micro Devices was started at $600, an "Agentic AI Sweetspot" with earnings forecast to grow at a 65% compound annual rate through 2030, while Marvell Technology was initiated at $270, citing a $120 billion Google agreement and an Oct. 6 analyst day as a potential catalyst. Arm Holdings rounded out the group at $320, with O'Connor highlighting its CPU intellectual property dominance and expansion into custom silicon and accelerator IP that he said could double earnings.
Investing.com·8dRead more →
Inference-Optimized Silicon4impact 4

DOJ probes whether Nvidia structured $20B Groq licensing deal to dodge antitrust scrutiny

The Justice Department is investigating whether Nvidia structured its $20B licensing deal with AI chip startup Groq to avoid antitrust scrutiny, Bloomberg reported, citing people familiar with the confidential inquiry. The investigation centers on a deal announced in December under which Nvidia acquired rights to Groq's technology while leaving the startup as an independent company, and as part of the transaction Groq CEO Jonathan Ross and COO Sunny Madra joined Nvidia. The arrangement has drawn criticism from US lawmakers, who have characterized it as effectively a takeover that could weaken competition in the AI-chip market. Regulatory questioning into the Nvidia-Groq transaction began earlier this year as part of the Justice Department's broader antitrust investigation, the people said, and the licensing-deal probe, first reported by The New York Times, could ultimately end without enforcement action. Groq did not submit the deal for antitrust review.
Seeking Alpha·8dRead more →
Inference-Optimized Siliconimpact 4

OpenAI Used Its Own AI to Design Jalapeno Chip

OpenAI CFO Sarah Friar confirmed at Goldman Sachs' Communacopia conference on September 8 that the company used its own frontier AI models to design the Jalapeno custom inference chip, marking the first confirmed instance of an AI lab using its own intelligence to architect its silicon. The chip, co-developed with Broadcom and announced in June 2026, achieved a nine-month design-to-tape-out cycle, the fastest ever reported for high-performance hardware, though the figure is company-claimed. This move is tied to OpenAI's aggressive pricing, including an 80% reduction for the Luna model to $0.20 per million input tokens and $1.20 per million output tokens, which has driven a 10x usage increase and supports 32% enterprise revenue growth from June to July 2026. By controlling the silicon stack, OpenAI aims to decouple margins from third-party cloud providers, reinforcing its position as a full-stack utility ahead of its confidential S-1 filing with an $852B post-money valuation. Broadcom CEO Hock Tan indicated small prototypes are expected in late 2026, with full-scale ramp slated for the first half of 2028, a period that will test the scalability of this recursive design model.
Yahoo Finance·9dRead more →
Inference-Optimized Silicon2impact 4

Qualcomm, Intel, and Chip Stocks Surge on AWS Deal and Manufacturing Milestones

Qualcomm, Intel, Amkor, Nova, and FormFactor shares traded up in the afternoon session after Qualcomm announced a multi-generational product collaboration with Amazon Web Services to develop custom AI data center infrastructure, while ASML, TSMC, and Intel achieved key milestones in next-generation chip manufacturing. Qualcomm's partnership with AWS focuses on co-designing successive generations of custom AI chips for high-performance inference and data center workloads, aiming to strengthen AWS's internal hardware ecosystem, reduce operating costs, and improve energy efficiency. Simultaneously, ASML and TSMC announced a joint initiative to transition from traditional 6-inch photomasks to larger 12-inch formats, targeting a pilot line by 2031 and full production by 2033, which will expand the print field and reduce manufacturing costs. Additionally, Intel Foundry and ASML confirmed that Intel has processed over one million wafers on its Intel 18A node using High NA EUV equipment. These coordinated advancements boosted investor confidence in sub-2-nanometer chip designs and next-generation AI infrastructure. Among the stocks, Intel jumped 9.5%, Amkor and FormFactor each rose 6.3%, Nova gained 3.7%, and Qualcomm advanced 3.3%.
Yahoo Finance·10dRead more →
Inference-Optimized Silicon

AMD Unveils Desktop AI Workstation Running 300-Billion-Parameter Models

AMD has introduced a new workstation-class system, the Ryzen AI Halo platform powered by the Gorgon Halo processor, which can run AI models with up to 300 billion parameters locally on 192 gigabytes of unified memory, eliminating the need for cloud computing. CEO Lisa Su announced the capability on the August 4, 2026 earnings call, extending the previous Ryzen AI Max+ and developer platform that supported models up to 200 billion parameters. The move targets on-premise buyers such as defense contractors, hospitals, banks, and sovereign research labs that prioritize data privacy and latency. However, the near-term revenue impact is expected to be minimal compared to AMD's data center business, which generated $6.72 billion in Q2, up 107% year over year, and represents 58% of total revenue. AMD's stock closed at $477.57 on September 4, 2026, up 4.69%, and has gained 123% year to date.
24/7 Wall St.·10dRead more →
Inference-Optimized Silicon

SEMIFIVE Begins Mass Production of HyperAccel's AI Accelerator on Samsung 4nm

SEMIFIVE, a global provider of custom AI semiconductor solutions, has commenced mass production of HyperAccel's data center AI inference accelerator, 'Bertha', using Samsung Foundry's 4nm process node. This marks SEMIFIVE's first large-scale production on that node and involves a 'Big Die' chip exceeding 500 mm², with SEMIFIVE delivering a complete turnkey solution from design to volume manufacturing. The company expects production volumes to grow as HyperAccel expands its services, and this program adds to SEMIFIVE's portfolio, which already includes security and HPC chips. SEMIFIVE secured KRW 42.3 billion in new mass-production orders in the first half of this year, nearly double its full-year intake of KRW 21.2 billion last year, with quarterly orders surging 71% from KRW 15.6 billion in Q1 to KRW 26.7 billion in Q2, and overseas orders accounting for 45% of Q2 bookings.
PR Newswire·10dRead more →
Inference-Optimized Silicon

AMD Unveils $100,000-Plus AI Workstation for 2027

AMD shares rose 2% in Asian markets after the chipmaker introduced a new workstation aimed at running very large AI models without relying entirely on cloud infrastructure. The Threadripper Halo Station combines AMD's 96-core Ryzen Threadripper PRO 9995WX processor with two liquid-cooled Instinct MI350P accelerators, expandable to four, offering between 288 GB and 576 GB of HBM3E memory. Designed for AI development and inference with models up to one trillion parameters, the system provides up to 16 TB per second of pooled memory bandwidth and a closed-loop cooling system managing roughly 1,550 watts of thermal output. Targeting corporate AI teams, researchers, and developers handling sensitive workloads locally, it places AMD in competition with Nvidia's DGX Station systems. AMD has not disclosed pricing or a specific launch date, though commercial availability is expected in 2027.
GuruFocus·11dRead more →
Inference-Optimized Siliconimpact 4

NVIDIA's $20 Billion Groq Bet Goes Live This Year

NVIDIA Corporation's Groq 3 LPX rack is in full production and will be deployed alongside its Vera CPUs and Rubin GPUs at neocloud Nebius later this year, according to Nvidia senior director Dion Harris. The rack marks the commercialization of technology from Nvidia's $20 billion purchase of Groq's assets in December, its largest deal on record. Each liquid-cooled rack packages 256 Groq chips and can deliver 3,400 tokens per second, a benchmark from Artificial Analysis. Groq's chips are manufactured by Samsung, unlike Nvidia's own GPUs, which are made by Taiwan Semiconductor Manufacturing Co. The rollout currently relies on Nebius as the only confirmed customer, and Nvidia has also committed up to $100 billion to OpenAI and $5 billion to Intel, among other investments.
CNBC·12dRead more →
Inference-Optimized Siliconimpact 4

Cerebras Backlog Hits $25.4 Billion, Largely Driven by OpenAI Deal

Cerebras Systems reported a $25.4 billion backlog of contracted work as of June, with a significant portion tied to a single OpenAI agreement. The AI computing specialist's adjusted revenue grew 103% year over year to $209.9 million in the second quarter, and management raised its full-year outlook to $880 million to $890 million. The backlog, which is nearly 29 times expected 2026 revenue, stems largely from a December 2025 master agreement where OpenAI committed to purchase 750 megawatts of computing capacity, valued at over $20 billion, with an option for an additional 1.25 gigawatts. Cerebras recognized $56.8 million in revenue from the OpenAI arrangement in the second quarter, about 32% of its GAAP revenue. The company expects to convert only about 22% of the backlog over the next 24 months, with most scheduled after mid-2028, and it is building capacity, including a new 165-megawatt data center in Finland. OpenAI has also advanced a $1 billion working capital loan to support construction. Despite the backlog, Cerebras shares trade at about $215, down 44% from their 52-week high, valuing the company near $51 billion, roughly 58 times expected adjusted revenue this year.
The Motley Fool·13dRead more →
Inference-Optimized Silicon

Coatue Opens Positions in Intel and Cerebras, Signaling AI Chip Trade Broadening

Coatue Management's second-quarter filing revealed new positions in Intel and Cerebras, suggesting the AI chip trade may be broadening beyond NVIDIA. The firm reported 12,084,027 Intel shares and roughly 7.01 million Cerebras shares as of June 30, with the Cerebras stake valued at about $1.5 billion. Intel's appeal lies in its strategic relevance, including its CPU franchise, foundry ambitions, and domestic manufacturing, while Cerebras offers a radically different wafer-scale architecture with Q2 non-GAAP core revenue of $209.9 million, up 103% year over year. However, Intel faces capital intensity and competitive execution challenges, and Cerebras's total Q2 GAAP revenue was $180.1 million with hardware sales down 23% to $54.1 million. The filing supports a broadening thesis but is not a verdict, as quarter-end holdings can change and subsequent 13F filings will test conviction.
Insider Monkey·14dRead more →
Inference-Optimized Silicon3impact 4

DeepSeek to Use 160,000 Huawei Chips in Mongolia Data Center

DeepSeek, a leading Chinese artificial intelligence company, plans to install at least 160,000 Huawei Ascend 950DT chips in a large data center under construction in Inner Mongolia. This could become one of the largest disclosed clusters of Huawei AI chips and marks a significant step for China in reducing its reliance on Nvidia chips. The chips will primarily be used for running AI models, or inference, while model training will still use Nvidia chips. However, Huawei's production capacity is insufficient due to a shortage of key components, especially high-bandwidth memory, which may delay delivery by more than a year. The project has a power capacity of 1 gigawatt, and the 160,000 chips will be only part of the total computing power. This cluster is 16 times larger than previously disclosed Ascend clusters, reflecting China's progress in building AI infrastructure with domestic technology amid competition with the United States.
Money & Banking·14dRead more →
Inference-Optimized Silicon6

CGSI Upgrades Electronics Sector to Overweight, Highlights DELTA as Top Pick

CGSI, or CGS International Securities (Thailand), has upgraded its recommendation for the electronics sector to Overweight from Neutral, and also upgraded DELTA to Buy with a target price of 360 baht, from Sell. The firm believes the transition to 800VDC technology will create business opportunities of approximately 20 billion US dollars. Meanwhile, the global market for power systems and cooling systems for data centers may expand from 50 billion US dollars in 2025 to 117 billion US dollars by 2030, representing an 18% CAGR. DELTA will benefit from its power management and liquid cooling products. HANA has upside from its power semiconductor business, especially silicon carbide (SiC) chips, through its subsidiary Powermaster, which will start generating commercial revenue in the second half of 2027. CGSI recommends Buy on HANA with a target price of 62 baht, and expects DELTA's revenue to grow 40% and EPS to grow 67% in 2027.
HoonSmart·14dRead more →
Inference-Optimized Silicon

ASUS Unveils Integrated AI Factory Platform at AI Tech 2026

At ASUS AI Tech 2026 in Seoul, ASUS unveiled a unified AI factory platform that integrates accelerated computing, networking, storage, deployment software, AI operations, and governance, partnering with NVIDIA and other technology firms including Aleria, AMD, Crusoe, Foxlink, IBM, Intel, Samsung, Schneider Electric, and WD. The platform aims to help enterprises shorten time to first token, accelerate time to revenue, and optimize infrastructure efficiency while maintaining trusted AI at scale. ASUS also introduced new hardware, such as the ASUS AI POD XA VR721-E3 built on NVIDIA Vera Rubin NVL72, which delivers 10 times the performance per watt versus the prior generation, and the enterprise-ready XA NR1I-E12LR and XA NR1I-E12L based on NVIDIA HGX Rubin NVL8. For agentic AI, the XA P2N-E2 supports up to two dual-slot NVIDIA GPUs, while the ESC8000-E12P and PE3000N extend acceleration to edge and visual computing. The company emphasized that AI governance starts before the first token, using simulation and digital twin environments to validate deployments, and extends through operations with tools like the ASUS AI Software Stack and Turnkey AI Applications. ASUS Senior Vice President Paul Ju stated that infrastructure performance and AI governance can no longer be addressed separately, as AI factories evolve into always-on production environments.
PR Newswire·15dRead more →
Inference-Optimized Siliconimpact 4

AMD Shares Down 5.2% Since Q2 Earnings Beat

Advanced Micro Devices shares have fallen 5.2% since its second-quarter earnings report, underperforming the S&P 500. The company reported non-GAAP earnings of $1.66 per share, up 246% year over year and beating estimates by 3.1%, while revenues rose 50.1% to $11.54 billion. Data Center segment revenues surged 107.3% to $6.72 billion, contributing 58% of total revenues, driven by EPYC and Instinct GPU demand. For the third quarter, AMD expects revenues of approximately $13 billion, plus or minus $300 million, implying 41% year-over-year growth. Management also announced that Anthropic plans to deploy up to 2 gigawatts of MI450-Series GPUs, with first shipments of its Helios AI platform expected in the third quarter.
Zacks Investment Research·15dRead more →