Megatrend · Robotics & Physical AI
We've taught robots to move — what's left is teaching them to "see" and "feel"
Motors and gears do make a robot arm move. But a robot that moves while seeing nothing and feeling nothing is like a blind person with both hands gone numb — it can move, but it can't do a single delicate task. This lesson is about the two senses that make a robot actually useful: the "eyes" (machine vision — the cameras and brain that inspect, locate, and guide) and "touch" (force and tactile sensors that let it pick up an egg without crushing it). This is the segment where Japanese companies earn god-tier profits, and the last gate before robots can grip and handle as well as humans.
01What it is (the robot's eyes and hands)
Picture someone very strong, able to lift heavy things, with arms that move in every direction — but with their eyes closed and their hands numb. Could they still do anything useful? Almost nothing. Because every real task, from putting parts in a box to assembling a car, requires "seeing where things are" and "feeling whether you're gripping tight enough." A robot is exactly the same. Motors and gears give it muscle, but this node is what gives it eyes and a sense of touch.
This node bundles two abilities that look unrelated but are really a pair:
- Machine Vision — the "eyes": high-resolution industrial cameras plus a brain that processes the image, doing the work of inspecting whether a part has defects · locating where an object sits and at what angle · and guiding the arm to reach in and grab the right spot. This is seeing to "inspect and control work" in a factory — a different thing from seeing 3D images to drive a car
- Force/Tactile Sensing — "touch": sensors that measure the force the robot pushes or pulls with, and the contact at its fingertips, so it knows whether it's gripping too hard, too lightly, or has bumped into something — the key to delicate handling
On the megatrend map, this node is a sub-branch under Robotics Components & Actuation within the big trend Robotics & Physical AI. It sits in the "upstream layer" — a supplier that makes the sensors and cameras that robot builders assemble into their machines. Its siblings next door are the reducer and the servo motor that act as the "muscles," plus LiDAR & 3D Perception, which makes the 3D distance sensors that let self-driving cars see the world around them — while this node focuses on the "eyes in the factory" (seeing to inspect and control work), not the "eyes on the road" (seeing to drive).
Machine Vision = a system that uses cameras + software to "see" and then decide in place of a person (this part passes/fails, where an object's coordinates are) · Force/Torque sensor (6-axis force sensor) = a device mounted on the robot's wrist that measures push, pull, and twist forces in all 6 directions at once, so the robot knows how hard it's pressing · Tactile sensor = a sensor at the fingertip or surface that "feels" the point of contact and the pressure at each specific spot, much like a human fingertip
02Why it matters — the gateway to precise handling
There's a saying in robotics: "moving is easy, grasping is hard." Today we can build arms that move fast and precisely without breaking a sweat. But almost everything robots still can't do is "hand" work — picking objects scattered loosely in a box, plugging a cable into a socket, grabbing something soft without wrecking it. They can't do these not because the arm isn't strong enough, but because the robot can't see finely enough, and can't feel. That makes this node the "gateway" robots must pass before they can do truly valuable work.
Start with machine vision — this market isn't small. The global machine vision equipment market was worth about $15.8 billion in 2025 and is expected to grow to roughly $23.6 billion by 2030, at a compound annual growth rate (CAGR) of about 8.3%. It's this big because it hides everywhere: electronics lines use it to spot defects on chips, car factories use it to guide welding arms, warehouses use it to read barcodes and locate items — every time a factory wants to cut down on people inspecting by eye, machine vision is the answer.
But the side that's about to explode is touch. Industrial robots used to need almost no force sensors, because they just repeated the same motion in a steel cage. But once the world started building humanoid robots that work in the real world alongside people, the story flipped. The market for force (torque) sensors made specifically for humanoids was worth about $515 million in 2024, but is expected to surge to billions of dollars by the early 2030s as more robots get built — because every humanoid needs dozens of force and tactile sensors.
Put simply — as long as robots have to do work that requires "seeing" and "gripping" (and that's nearly every job worth doing), this node is an indispensable layer. It holds both a slow-but-steady business (factory vision) and a business about to grow exponentially (touch in humanoids) in one.
03How it works (see → think → tell the robot)
The heart of machine vision is a short loop that spins very fast — fractions of a second per part. Let's walk through, step by step, how "light hitting an object" turns into "the robot knowing what to do."
The "hard" parts of vision are steps 1 and 3. Step 1 — light and lenses — sounds mundane but matters enormously. If the lighting is bad, the defect you need to see vanishes into the shadows. Half of a machine vision engineer's job is "arranging the light so the thing you want to see stands out." Step 3 — processing — is the battlefield changing fastest, because it used to require writing fixed rules ("if a black spot is bigger than X = defect"), but now uses deep learning that learns from sample images, making it far better at inspecting objects whose appearance varies (like scratches on a natural surface).
Touch works in a completely different way. It doesn't "look first, then act" — it's real-time feedback (the orange loop in the figure). While the robot is gripping something, the force sensor at the wrist or the tactile sensor at the fingertip constantly measures how hard it's pressing. If it grips too hard and starts to break the object, it eases off; if it's too loose and the object is about to slip, it adds force — just like when we pick up a glass of water, we don't calculate the force, we adjust constantly to the feeling at our fingertips.
04Where it sits in the robotics world
If you think of a single robot as a body, the motor and gears are the muscles and joints, and this node is the eyes and senses. The two have to work in concert — strong but blind muscles can't do delicate work, and good eyes with twitchy muscles can't grab accurately either.
- Always paired with the motor/servo: vision says "the object is at these coordinates, tilted 15 degrees," and then the motor moves to grab it. While gripping, the force sensor tells the motor to ease off when it touches the object — these three are the closed loop of handling
- Different from its sibling LiDAR & 3D Perception: both are "vision," but for different jobs — LiDAR fires lasers to measure distance and build a 3D map, so cars and mobile robots know "how far away things are," while the machine vision in this node looks at high-resolution 2D images to "inspect and control work within arm's reach." Different problem, different players
- The brain comes from AI: the smarter the AI model, the more accurate the vision and the reading of touch become — deep learning is what lets modern vision detect objects it once couldn't, and turns tactile data into smarter decisions
- Uses chips from Semiconductors directly: the heart of a camera is the image sensor, which is a kind of chip, and image processing needs processing chips — so this node is a real customer of the semiconductor industry. Whoever controls the image sensor controls the upstream of the entire camera world
- Feeds Industrial Automation and robots of every kind: the real customers are every factory that wants to inspect by machine instead of by hand, and every humanoid being built
05Where it stands now
The most astonishing thing about the machine vision side is that it's an unbelievably profitable business. The market is led by two players — Keyence (Japan) and Cognex (U.S.) — who together hold about a quarter of the global machine vision market (Keyence ~14% + Cognex ~11%). Keyence is a legendary case study: in the fiscal year ending March 2025 it broke ¥1 trillion in revenue for the first time, with an operating margin of about 52% and a gross margin around 80% — numbers a typical manufacturer can only dream of (usually below 10%). The secret is a "fabless + direct-sales" model: it owns no factories of its own, but sends engineers right into the customer's site to solve problems, then sells premium-priced sensors.
The difference between the two is interesting. Cognex is "pure machine vision" — vision is its core business, strong on flexibility and a deep-learning toolkit (ViDi) that lets customers tune it deeply themselves. Keyence, meanwhile, sells a full range of sensors (vision being one), and wins on "ease of use + on-site service + high margins." Meanwhile Teledyne (owner of DALSA and FLIR thermal cameras) and Sony control a key upstream — the image sensor — with Sony holding about 43% of the global image sensor market in 2025. Most high-end machine vision cameras are likely using a Sony sensor inside.
The other side is China catching up on the camera front — Hikvision and Sunny Optical rose from the surveillance-camera and phone-lens markets, then moved into the industrial machine vision market on the strength of price, pushing mid-grade products ever cheaper.
The touch (force/tactile) side is still a smaller and more fragmented market. The standout players are specialist firms, many of them still private — for example ATI Industrial Automation (the leader in 6-axis force sensors for robot wrists), FUTEK, and Bota Systems, fighting for share in the wrists of collaborative robots (cobots) — and this is the side the humanoid wave is about to change. Because robots that work alongside people need to "feel" for safety and precision.
06What's ahead — deep learning, e-skin, humanoids
The first direction shaking up the vision side is AI swallowing traditional machine vision. Inspecting by image used to require engineers to write rules one by one, but deep learning changes it to "feed the machine examples of good and bad images and let it learn for itself," making it much better at inspecting objects whose appearance varies. Analysts expect AI-integrated vision systems to make up over 55% of new industrial installations by 2033 — those slow to adapt to AI will lose ground.
The second direction is the touch side about to grow from a niche market into a mass market. The numbers tell the story clearly: the force-sensor market for humanoids is expected to grow from ~$515 million (2024) to billions of dollars in the early 2030s, because every humanoid needs dozens of force and tactile sensors. The more robots get built, the more demand grows exponentially.
The third direction is e-skin (electronic skin). Today most robots can still feel only at a few points on their fingertips, but research is trying to build "artificial skin" that senses contact across the whole hand and arm, like human skin with receptors spread throughout. If it can be done at a price that's actually manufacturable, robots will be able to grip soft and fragile objects and do delicate assembly work at a level that isn't possible today — and this is the "last gate" to making a robot's hands as skilled as a human's.
The fourth direction is vision and touch fusing together. Modern robots don't "look then act" or "feel then adjust" separately — they use both at once, glancing roughly at where the object is, then using touch to fine-tune as they actually grip, just like people do. Combining the data from eyes and hands (visuo-tactile) is the hottest research direction in teaching robots to grasp.
07Challenges & risks
The first risk is that AI could become both friend and foe. Deep learning does make vision better, but it also erodes the "specialist know-how" that used to be the market leaders' wall. Setting up a vision system used to require experts who could write rules well, but if one day anyone can drag in images and train a ready-made model, the value could shift from hardware sellers to software and new software players — incumbents have to keep up or get disrupted.
The second risk is commoditization, especially on the camera and mid-grade sensor side, where Chinese makers like Hikvision and Sunny Optical come in and undercut prices. Products that used to sell high because they were "hard to make" now have rivals at half the price, squeezing incumbents' margins — even though the very top tier (like Keyence) can still hold premium pricing, the mid-market is a price battlefield.
The third risk is the industry's cycles. Machine vision is tied to factory capital spending, especially electronics and auto factories, which rise and fall in cycles. When the economy slows, factories delay investing in new machines, and vision sales soften along with it — 2024–2025 was a stretch where the vision market went quiet for a while before recovering.
The fourth risk is on the touch side, which is still a "promise of the future" more than real revenue today. The tactile-sensor market for humanoids is still small and rests on the assumption that humanoids will be made by the millions. If humanoids come slower than hoped (and robotics tech is always "slower than promised"), the big demand that's been projected gets pushed back.
In short: in an era where everyone talks about the "brain" and "muscles" of robots, what gets overlooked is the "eyes" and "fingertips" — the two senses that separate a robot that merely moves from one that does the work. Understanding this node fully is understanding why Japanese sensor companies that ordinary people have never heard of can earn profits several times higher than far more famous tech companies.