This Week in Robotic Learning (TWIRL) #8: How robots sense the world
And how they don't. And how they could.
When autonomous robots enter our lives, one obstacle will be that in many respects they will think and act like us, but they will not sense the world as we do. A human has five senses: vision, hearing, touch, taste and smell. A robot will certainly have a version of some of those that is well beyond human capacity and lack others altogether. It may have senses that are not found in nature let alone in humans.
There are a wide range of “senses” that we simply do not have. I, for one, have never walked into a room and thought it seemed a bit more radioactive than usual or into a cave and used echolocation to find out where all my fellow bat bros are hanging out (pun intended, but immediately regretted).
The strength of our senses–and of our robots’ senses–does not need to be perfect. We cannot smell as well as a dog, nor see as well as a hawk, nor, evidently, taste as well as a catfish. And yet, we get by. The goal isn’t to match humans, it’s to be sufficient to the task.
As these robots are in our homes, our stores and our lives, we will not be able to fully resist the urge to anthropomorphize them (and given their autonomous intelligence, may not be wrong to do so). Trying to categorize their ways of sensing the world within the framework of human senses is a useful way of trying to understand what these things are, how they’re like us, and how they’re not.
We’re still defining what we want our robots to be. It’s worthwhile to consider which senses they should have and to what degree of precision is necessary, and in what cases. This piece is mostly going to focus on the hardware side rather than the software side, so more “can the robot see?” rather than “can the robot intelligently interpret what it’s seeing?”
Of the five human senses, the vast, vast majority of the focus to date has been on vision. However vision on its own is insufficient. There’s more to evaluating the world than cameras. Consider this a guide to what’s possible now and may be in the future for the robotic senses.
Vision
It’s cameras. You can keep piling cameras on a robot until it sees everything, everywhere, all of the time.
3D camera manufacturer RealSense, which spun out of Intel in 2025, last year claimed to be integrated into nearly 60% of humanoid and autonomous mobile robots, so their stuff should be regarded as pretty typical of what’s available on the market. These are pretty similar to the camera in your phone, but have a bit better depth perception.
The innovation over time in robotic vision has been more about adding cameras to places where people don’t have eyeballs–the wrists/hands, and head/torso. Look at Sunday’s Memo, which has, for example, hand cameras, and wears a cute li’l hat to have a camera facing straight down that you could imagine is useful for, say, chopping vegetables. I could see how “more cameras” and “more degrees of freedom in limbs” could open up a new morphology that is basically a stick on wheels with more flexible arms like General Grievous from Star Wars.
The future of robotics??? Probably not. But maybe!
Vision is, if not a fully solved problem in robotics, one where improvements are “nice to haves” at best. The cameras are good at seeing whatever is available.
One question that’s interesting with vision, but also hearing and touch is: are we measuring the thing or changes to the thing? In all cases, measuring the thing itself is more thorough, but costlier in terms of expense, latency and battery life.
For vision, measuring the thing is using a typical RGB camera. Measuring changes to the thing entails using what’s called an event cameras (aka neuromorphic cameras). Rather than capturing the full image like a standard camera, they capture changes to the image.
You cannot watch a video produced by an event camera. Their output is data. Event cameras don’t seem super likely to be directly used in general purpose robots in the near future, but they might represent a pathway to cheaper data collection, and if used directly in the bots, a better battery life. Not super likely to occur any time soon, but you never know.
To the extent there’s value to be gained with robotic vision, it’s more a data compression/data creation problem (being pursued by academics rather than industry) than a “can a robot see stuff?” problem. It can very much see stuff. More than we can. More than is necessary to the task.
Hearing
It’s microphones. This one’s pretty much solved too. Here’s one for $55.
As anyone who has ever worked as a police informant knows, microphones do need to at least be at least somewhat oriented toward the source of the noise, and as anyone who has ever tried to have Alexa tell them the weather knows, there can be some shakiness in the handoff from the sound recording to the AI system, particularly if there are any other competing noises in the room.
Like voice assistants, modern robotics audio is optimized for voice commands, rather than making direct use from ambient noise (you can sneak up on a robot and it won’t suspect a thing if it doesn’t see you). But that doesn’t necessarily have to be the case. I came across an interesting paper recently from some authors at East China Normal University (go Lions!) and Fudan University (mascot unclear!), adding to the various Vision-Language-Action (VLA) model families I’ve reviewed in this space by building in some audio sensors to get more learnings from their surroundings.
Their research involved building in their own audio to a simulated environment, which isn’t exactly grading your own homework, but it’s also very much not succeeding in the real world. To be clear, this research line is very early days.
Robotic hearing is already sufficient for the necessary task of taking voice commands. But the endgame for this research, should it bear out, is pretty appealing: audio in robotics isn’t like audio for humans–it’s a direct supplement to touch. Gluing on cheap microphones near contact points could be helpful for figuring out, say, if an object is solid or not, especially in environments with non-ideal lighting.
Touch
Touch is the most compelling of these in that it is, currently, the worst relative to the value add it could be providing.
I do want to be clear: when I say “touch” I pretty much exclusively mean “tactile sensing.” What I do not mean is proprioception, that is, a robot’s sense of where its body is at a given moment. That’s super important too, of course, but not really a “sense” in the classic meaning.
Setting proprioception aside, within robotics today, touch can be divided into two categories: force/torque sensing (how hard and in which direction the force is applied) and tactile sensing, which covers pretty much everything other than an object’s weight: where it is, its shape, edges, texture, the onset of slip, temperature and vibration. Force/torque sensing is pretty evolved at this point and has been around for many decades. Tactile sensing is a bigger problem, and thus represents a bigger opportunity.
I’ve discussed data quality and quantity bottlenecks repeatedly in this space, as visual data (videos, simulations, etc.) is much sparser than the textual data LLMs trained on. It’s bad for visual data–for touch data it’s much worse.
Touch data requires… touching stuff, which involves wearing out the device that does all the stuff-touching. Touch data is not really transferable from device-to-device. One camera’s video is more or less the same as another–not true for tactile sensors. Video-type data can be gathered in simulated environments–not so much for tactile data (maybe world models solve this?).
Still, this is a really nascent space, so it’s possible that this field scales quickly. Sunday and Genesis are clearly betting on first party “gloved” data collection to improve the sense of touch.
Tactile sensors in a robot don’t work quite the same way that human touch does. Transduction, or the process of converting a physical touch into an electric signal, comes in a bunch of flavors; these are the most interesting ones in modern robotics. The precision you need depends in whether the goal is merely avoiding a collision or dextrous tasks–from least to most dextrous:
Capacitive - Two thin conductive layers separated by a squishy gap, when pressed closer together changes a measurable electrical property called capacitance. For robots, this is useful for large-area “skin” and collision-detection safety. It’s not going to solve grip-level touch, though. (Ex: your phone screen and all touch screens).
Magnetic - Put some magnets under a rubber skin. It’s thinner and cheaper than optical, but less performant and subject to magnetic interference. (Ex: ReSkin)
Vision-based / optical tactile - A small camera inside a finger, looking outward at a soft, opaque gel pad that is lit from the side. When the gel presses against an object, the camera films the resulting dent in extreme detail. On the plus side, this creates spatial resolution (how finely a sensor can tell apart two separate points of contact) with higher precision than a human fingertip. On the downside, it’s bulky, adds latency and the gel wears out. Despite those downsides, the upsides make this the dominant sensor tech for research and appears to be a promising path forward. (Ex: GelSight)
As with video (event cameras v. standard camera) and using audio as a touch substitute, there still isn’t clarity on how capturing the thing (i.e the touch) and capturing the derivative of the thing (i.e. what has changed) will play together. And as with video, capturing everything is more expensive and battery-intensive than only capturing changes. As such, the potential upside of relying only on sensing changes to the “touch picture” rather than the whole thing all the time is lower latency and lower power usage.
GelSight is the clear market leader in optical tactile tech, however I’m not sure if any robotics company at the frontier has integrated an off-the-shelf tactile solution. Figure, Tesla, Sanctuary AI are building humanoids with custom tactile sensors.1x, Apptronik and Boston Dynamics may not have invested in tactile sensing at all.
Among companies specifically targeting the touch problem: Eka Robotics has built a Vision-Force-Action model focused more generally on force-based sensitivity rather than tactile sensing (though these types of touch aren’t totally discrete–they too have built their own tactile grippers).
As mentioned earlier, this type of data isn’t really transferable from device-to-device and there’s not very much of it. To deal with that, Genesis AI is building a glove a person wears that corresponds to a hardware hand. Maybe it’ll work, maybe it’ll be “there are now 15 standards.”
Meta Digit Plexus
Meta is taking on the same problem from the other side, attempting to build an open infrastructure stack consisting of DIGIT (a fingertip sensor), DIGIT 360 (a really fancy fingertip sensor that can also sense vibration and heat) DIGIT Plexus (a standardized hardware surface that can have the DIGITs and ReSkin mounted on it, in an effort to start solving the integration problem) and Sparsh (and Sparsh-X), a tactile foundation.
So, this is all well and good. But hopefully it’s clear that at this point, touch is at most, a condiment, not the main course. Most robots rely on vision and proprioception, full stop. So the question is not “what does touch do today?” so much as what it can do tomorrow. And within that, there are some open questions.
Is force plus a good model sufficient or do we need rich, tactile sensors? Will durable skin become widespread, or will it be metal exteriors/sweaters from here on out? Will building a VLA-type model with touch integrated from the jump produce superior results? Can we all agree on a standard?
Do we even need touch or can we just keep piling on cameras?
I think we do. Yes, touch is hard. Physical Intelligence Co-Founder Sergey Levine has said he expects changing a baby’s diaper to be among the last tasks a robot masters. But it’d be really nice if a robot could change a baby’s diapers!
The theoretical possibilities of “really, really, good touch” well beyond the visible frontier are extraordinarily compelling. Human touch stops at the skin of the object; robots could go deeper, sensing which peach is bruised, or if there is in fact, gold in them thar hills. It could understand your health and mental state from a handshake. It could sense changes in your physical strength over time that would be too subtle to notice as they’re happening. As it felt you break down, it could sense itself breaking down too–and repair you both. It could tap on your head the way a plumber taps on a pipe and find a tumor. It could move from being a single sense to a universal measurement interface for the physical world.
That all seems pretty cool, so get on it, nerds!
Taste
What am I, a medieval king worried that a scheming courtier might poison my soup?
Robots do not need taste. Humans barely do at this point. We’ve written down which foods are healthy–thanks, natural selection, but we can take it from here.
That said, if innovations in taste sensors lead our best food scientists to surpass their most recent important breakthrough innovation (by which I mean Nerds Gummy Clusters–shout out, Mike Leach) then I take it all back.
Smell
See: Taste. I don’t really need my robot to have this except in very rare, niche cases. NEXT.
Non-human senses
There are a range of senses that don’t really have a human equivalent, but could be integrated into a general purpose robot.
Lidar sweeps a laser across a field of vision and times each reflection. It could be used to supplement vision (and is, in some autonomous vehicles) by providing a geometrically accurate 3D map, regardless of lighting conditions. In fact, it already is reportedly used by Unitree, Boston Dynamics and Agility.
Boston Dynamics’ Spot quadruped robot is basically a non-human sensor on four legs. It can be customized with thermal cameras (for overheating motors and electrical faults), optical gas-imaging sensors and gas detectors (for leaks), acoustic imagers (to “see” ultrasonic air leaks and mechanical faults), and radiation monitors (better a robot than me on that job).
Boston Dynamics Spot
There are many more senses that could be put on a robot that a human doesn’t have. Radar, magnetic field sensing and WiFi sensing can be backups and supplements to vision in various conditions. Ultrasonics could be used to improve performance on handling glass or clear objects. Thermal imaging could help robots interact more delicately with humans (and also dominate hide and seek). Ultraviolet sensors could make cleaning biological residues simpler and more accurate. Hyperspectral imaging–more granular cameras–could identify materials directly so that a farm bot could know what was worth harvesting. Terahertz imaging (like airport body scanners) and chemical trace detection (like the airport “swab” test) could be used for always-on, non-invasive security and general threat detection. GPS could be used to go on some fun hikes with your robot buddy.
The Spot quadruped provides an instructive example of how these sensors designed for niche use cases could be integrated into a modular chassis.
Wrapping up
Chart meant to be illustrative
Taken as a whole when we line each sense against what we want from it rather than against human capacity, the picture is clear. Robots are sufficiently adept at their vision and hearing, which is why the open questions are about data efficiency. On touch, though, demand beats supply. Even though robots can already beat a human fingertip on raw spatial resolution, it cannot change a diaper. That’s why I spent about half this thing talking about touch! It’s the biggest gap against the jobs to be done.
And while the current research is exciting, it’s still clearly pretty far behind what humans are capable of. Given the “hmmm… no great solutions on the horizon” state of affairs, alternatives like Eka’s Vision-Force-Action approach become more appealing.
Even with audio and visual, while the data intake is accurate, it isn’t necessarily what the model needs. That may be a model problem or it may be a data problem, but in either case, there may be benefits to moving away from the VLA framework toward a unified multi-sensor architecture that interprets more senses collectively.
OmniVLA is an early research stab at this, fusing an infrared camera, mmWave radar, and a microphone array on top of a standard RGB camera. (Despite my best efforts to focus on hardware here, some discussion of software is necessary.)
On the whole, this is encouraging! We’ve proven the ability to take huge quantities of data and drop it into a transformer and make it useful in LLMs, and to an extent, in robotics. There’s no reason that data must only be visual and audio. There’s a great opportunity to be creative here and it’ll be fun to see who captures it.
Let’s look at some robots
I’ve been somewhat critical of the humanoid form factor in the past, but I think this is a really cool example of a sort of “post-humanoid” design. It’s strong! I really also like the way it’s capable of rotating–that seems pretty clearly useful (though again, as per earlier in this piece, the head rotation at least could be obviated by more cameras. The face just being a big ring light is a nice way to bring your own lighting.
On each rewatch I am more and more confused by the directorial choice to have the robot bring the fridge to a guy who’s just, like, playing Candy Crush on his phone. Were they trying to show his nonchalance? Because he just looks bored, though he does perk up a bit when he gets to drink his off-brand sparking water. Sorry that the frontier of human innovation doesn’t interest you, buddy! (I have to assume this guy was integral to building this robot. But still!)






