AI for Manufacturing: What Actually Deploys
Cicero Campelo, CISSP
August 18, 2026 · 14 min read
Part of our guide to AI for startups.

Table of contents
- What AI for manufacturing actually means
- What are the main use cases for AI in manufacturing?
- Why guidance is the layer that ships first
- Why manufacturing needs AI: the labor math the use case lists skip
- What actually kills a factory floor pilot
- The camera problem: process IP, surveillance, and retention
- Where a small team still wins
- What to do this week
- Sources
- Frequently asked questions
AI for manufacturing means putting machine learning, and increasingly multimodal models that can process video and speech, to work in a plant in one of three ways: seeing what is happening, telling a person what to do about it, or acting on the process directly.
Nearly every page that ranks for this term hands you the same list. Predictive maintenance. Visual quality inspection. Demand forecasting. Generative design. Digital twins.
The list is accurate. It is also not the decision in front of a founder, because those use cases sit at different layers of the same stack, and the layers have wildly different deployment costs. One of them can realistically clear a plant's procurement and safety review in a single budget cycle. The others usually cannot.
In February 2026, Y Combinator published a short request for startups on AI guidance for physical work. Its framing is the one sentence that changes where a founder should enter: across field services, manufacturing and healthcare, AI cannot yet act in the world. What it can do, in YC's words, is "see, reason, and guide the human who does."
That single constraint reorders how you enter manufacturing.
What AI for manufacturing actually means
Strip the vendor categories away and there are three layers, in increasing order of difficulty.
Layer one: sense. The system observes and reports. Camera-based defect detection on a line, vibration analysis on a bearing, thermal anomaly detection on a press. The output is a flag or a score, and a human decides what to do about it. This layer has been shipping since well before the current model generation, and most of what gets marketed as AI in manufacturing lives here.
Layer two: guide. The system observes, reasons about the situation, and tells a specific person what to do next. YC's illustration is a worker wearing a small camera while a model watches what they see and talks them through the job: "Turn off that valve. Use the 1/2-in wrench. That part looks worn. Replace it." The human still does the work and still owns the decision. The model supplies the judgment that used to take years to acquire.
Layer three: act. The system closes the loop and moves something. Robotic cells, autonomous material handling, closed-loop process control that changes a setpoint without asking. This is the layer people mean when they say AI is coming for manufacturing jobs, and it has the longest deployment cycle by a wide margin.
Most writing about AI for manufacturing blurs layers one and three and skips layer two entirely. Layer two is the one that only recently became buildable, and it is the one a small team can actually get into a plant.
What are the main use cases for AI in manufacturing?
Here is the list every vendor page gives you, with the layer each one sits in and an honest read on how far along it is.
- Visual quality inspection. Layer one, and the most mature thing in the category. A camera over a fixed station is a solved data problem and the output is a pass or fail that a human can override. Entering here means competing with established machine vision vendors rather than finding a green field.
- Predictive maintenance. Layer one, and the most oversold. The modelling is not the hard part. Getting labelled failure history off equipment that was never instrumented is, and a plant often has not recorded enough failures of a given mode for anything to learn from.
- Process optimization. Layer one or layer three depending on one decision: whether the system recommends a setpoint or changes it. That single choice is the whole deployment problem in miniature.
- Demand and supply planning. Layer one, and it lives in the enterprise systems rather than on the floor, which is why it clears an IT review instead of an OT review. This sits closer to AI for supply chain than to anything happening on the line.
- Generative design and simulation. Layer one, and engineering rather than operations. It changes what gets made, not how a shift runs.
- Real-time work guidance. Layer two, and the newest entry on the list. A worker wears a camera, the model sees what they see, and it talks them through the job. This did not exist as a product category before models could reason about video.
- Robotic cells and autonomous material handling. Layer three. Real and deployed, but capital-intensive, which means a long sales cycle and a customer who is buying equipment rather than software.
Read that list as a founder instead of as a buyer and one thing stands out. Six of the seven are categories where large vendors have been selling for years. One of them is new.
Why guidance is the layer that ships first
YC's case for timing rests on three converging conditions, and each one is checkable.
Multimodal models can now see and reason about real-world situations reliably. Not perfectly, but reliably enough that a wrong answer is a conversation rather than a scrapped part.
The hardware already exists. YC's list is phones, AirPods, and smart glasses. Nobody has to buy capital equipment or requalify a line to run a pilot. That matters more than it sounds. A layer-three deployment starts with a purchase order for machinery and a requalification plan. A layer-two pilot starts with a phone on a lanyard.
And the economics are urgent. YC's own framing is that skilled labor shortages "make this economically urgent" and that the work involved is high-wage employment for millions of people.
There is a fourth reason YC does not spell out, and it is the one that decides pilots: at layer two, being wrong is cheap. A bad recommendation costs a conversation with a supervisor. A bad actuation costs a part, a line, or a person. That asymmetry is why the safety review that stretches into months for a robotic cell can be a meeting for an advisory app. The same sequencing shows up across the physical economy, from AI in agriculture to AI for field service: ship the thing that only has to see, then earn the right to act.
Why manufacturing needs AI: the labor math the use case lists skip
The reason a plant manager takes your meeting is not that the technology is interesting. It is that the person who knew how to run the machine retired.
Deloitte and The Manufacturing Institute project that US manufacturing could need as many as 3.8 million new employees between 2024 and 2033, and that roughly 1.9 million of those positions could go unfilled if the skills and applicant gaps are not closed. That is more than half the openings.
The headline number understates the problem, because the scarce input is not headcount. It is judgment. On an a16z panel, Turner Caldwell, who spent nearly a decade at Tesla before co-founding Mariana Minerals, describes the shape of it in refining. The feedstock varies constantly, so temperatures, flow rates, chemical addition rates and residence times need continuous tuning, and in his words, "we don't have that labor pool here that has that embedded know-how." He puts the mining industry at roughly 35 years of meaningful attrition in its labor pool.
Caldwell's own answer to that gap is autonomy, reinforcement learning that takes people out of the refinery control loop, which is a layer-three bet. He can make it because he owns the plant. Most manufacturers cannot make it this year, and the gap is there either way.
Refining is not discrete manufacturing, but the shape is identical. The operators who can walk up to a line and hear what is wrong are leaving, and the apprenticeship pipeline that produced them has thinned out.
That gap is exactly what layer two closes. YC's version: "Instead of needing months or years of training, workers can be effective immediately with AI coaching them and accessing new skills when needed."
It is also where the budget is. In a companion request for startups on new operating systems for the physical world, YC notes that "80% of the global workforce doesn't actually sit at a desk, but the software for the physical world hasn't really changed in over 20 years," and that these industries "spend 10 to 100 times more on labor than software." If your product moves a labor number instead of a software number, you are selling against a budget line one to two orders of magnitude larger. That is the service as software argument pointed at atoms.
What actually kills a factory floor pilot
Model quality is rarely why an AI for manufacturing pilot dies. These are the reasons it usually does.
The equipment has no data out. Plants routinely run machines from several decades and multiple vendors, many with proprietary protocols and no API. Getting a usable signal off an old press is a systems integration project, not a machine learning one. Layer two sidesteps this, which is another argument for entering there: a camera pointed at a person needs nothing from the machine.
The network will not let you out. Plant floors are segmented on purpose. The ISA/IEC 62443 standards for industrial automation and control systems formalize zones and conduits, with restricted data flow as a foundational requirement, and in practice that segmentation is usually laid over some version of the Purdue model. Direct calls to a cloud API from the control layer are not an oversight for you to route around, they are the design. If your architecture assumes a round trip to a frontier model from the floor, you will meet the operational technology (OT) security team and you will lose. Decide where inference runs, on device, on premise, or through a broker in the demilitarized zone, before the first customer call.
Downtime is not a variable you get to test. A running line has no staging environment, and the plant's scheduled downtime is spoken for months ahead. Your pilot has to run alongside production without touching it, which means read-only or advisory by construction, not by choice.
The approver is not the buyer. The person excited about your product is usually a continuous improvement lead or a plant manager. The people who can stop you are environmental health and safety (EHS), the OT security team, quality (because anything touching a controlled process may trigger requalification), and in many facilities a union or works council. Any one of them is a veto. Map all four in the first meeting instead of discovering them in month five.
The operator will not type. Gloves, noise, safety glasses, and no free hands. An interface that assumes a keyboard is dead on arrival. Voice and camera are not a nice touch in this market, they are the only viable input, which is a large part of why this became buildable only recently.
None of these are solved by a better model. They are solved by a team that spends time on the floor, which is the forward deployed engineer motion applied to a physical plant.
The camera problem: process IP, surveillance, and retention
Layer two of AI for manufacturing requires putting a camera on a worker inside a facility. Three consequences follow, and each of them has stalled real deals.
The plant's process is its intellectual property. A camera pointed at the line captures tooling, fixtures, sequences and settings that the company treats as trade secret. You are asking a manufacturer to stream its most valuable secret to a vendor's cloud, and often to a model provider sitting behind that vendor. Expect a legal review longer than your sales cycle unless you can answer four questions before they are asked: where inference runs, what leaves the site, what is retained and for how long, and whether their frames train anything.
The worker is a person, not a sensor. Continuous recording of an identifiable employee is surveillance, and in much of the world it triggers notice, consent, and negotiation. Germany's Works Constitution Act gives the works council a co-determination right over "the introduction and use of technical devices designed to monitor the behaviour or performance of the employees," which is a negotiation rather than a checkbox. Design for it: processing on the device, faces and bystanders blurred, a short local retention window, and an explicit written commitment that footage is not used for performance management. Bringing that to the first meeting turns a blocker into a differentiator.
Retention is the real liability. The instinct is to keep everything, because the recorded work is the data moat. Keep everything and you have also built a discovery target, a breach target, and a regulatory exposure that grows with every customer you add. The workable version separates the derived training signal from the raw frames, keeps the first, and expires the second on a short clock.
These are architecture decisions, not compliance chores to hand to someone later. They are cheap to design in and expensive to retrofit, which is the oldest lesson in security applied to a new surface. The same discipline that governs a data for AI strategy applies here, with the added weight of a camera pointed at a person.
Where a small team still wins
YC lays out three shapes for a company in this space, and they are genuinely different businesses.
Sell the guidance system to companies that already have a workforce. The fastest path to revenue and the hardest to defend. You are a tool, procurement will treat you like one, and your moat has to come from accumulated data rather than the interface.
Pick a vertical and build the full stack. YC's examples are HVAC repair and nursing. The manufacturing analogs are narrow on purpose: weld inspection, CNC setup, changeover on one class of packaging line. You own the workflow, the training content, and eventually the labor itself. Slower, far more defensible, and essentially the vertical SaaS pattern with a labor component attached.
Build the platform that lets anyone become a skilled worker. Someone signs up, the guidance system makes them employable in a trade they have never done, and the platform books the work. The largest outcome of the three and the hardest start, because it needs supply and demand at the same time, and because the first cohort has to be good enough that a customer takes a second booking.
Whichever shape you pick, the asset is the same one, and YC names it plainly in the companion video: a company built here gets to "record all the work as it actually happens," and in YC's framing neither the frontier models, nor robotics startups, nor today's software incumbents will have that kind of end-to-end data. That is the honest answer to what stops a foundation model from eating you. It is not your prompt and it is not your interface. It is a proprietary record of how the job is really done, in your vertical, at a level of detail nobody has ever captured. The competitive moat logic is the same as in software, except here the data does not exist anywhere else at all.
One test before you commit. Ask a prospective customer to name the number your product moves and who owns it. If the answer is overtime hours owned by the plant manager, first pass yield owned by quality, or unplanned downtime owned by maintenance, you have a buyer. If the answer is innovation, you have a pilot that gets quietly defunded in the next budget cycle. For where autonomy eventually lands and what the rest of this stack looks like, read physical AI.
What to do this week
- Pick one job on one line, not an industry. Changeover on the number three filler is a product. AI for manufacturing is a market report.
- Stand next to the person who does that job for a full shift. Write down every question they ask a more experienced colleague. That list is your first version.
- Draw the three layers for your idea and mark which one you are entering. If you are entering at act, write down what would have to be true to enter at guide instead.
- Ask your design partner for their network diagram and their OT security contact before you write a line of code. Let the answer decide where inference runs, not the other way round.
- Write a one-page data posture: what the camera sees, where it is processed, what is retained, for how long, and what it never trains. Bring it to the first meeting rather than the fifth.
- Name the number. Get the customer to say out loud which metric moves and who owns it, then get that person into the room.
Manufacturing does not reward a fast demo. It rewards the team that shows up on the floor, ships the layer where being wrong is cheap, and earns the right to act. For how this fits with the rest of what you are building, start with AI for startups.
If you want the operating system for running a company this way, that is what we are building in AI Operating System for Startups.
Sources
- AI Guidance for Physical Work, Y Combinator, February 2026: the request for startups this article distills, including the see, reason and guide framing, the three converging conditions, and the three company shapes.
- New Operating Systems for the Physical World, Y Combinator, July 2026: the deskless workforce figure, the labor versus software spend ratio, and the end-to-end data argument.
- The Founders Who Left Tesla to Rebuild America, a16z: Turner Caldwell on the embedded know-how missing from the industrial labor pool and on attrition in mining.
- Manufacturers Need as Many as 3.8 Million New Employees by 2033, The Manufacturing Institute, with the underlying Deloitte and Manufacturing Institute manufacturing skills gap study.
- ISA/IEC 62443 series of standards, International Society of Automation, on zones, conduits and restricted data flow in industrial control networks.
- Works Constitution Act, official English translation, section 87, on works council co-determination over monitoring technology.
- Profiles: Turner Caldwell and Mariana Minerals.
Frequently asked questions
What is AI for manufacturing?
AI for manufacturing is the use of machine learning, and increasingly multimodal models that can process video and speech, to do one of three jobs in a plant: sense what is happening, guide the person doing the work, or act on the process directly. Sensing covers visual defect detection, vibration analysis and anomaly detection, and it reports to a human who decides. Guiding means a model watches a job through a camera and tells a specific worker what to do next, while the human still performs the work. Acting means closed-loop control, robotic cells and autonomous material handling, where the system changes something without asking. The three layers are usually lumped together in vendor material, but they have very different deployment costs and timelines.
What are the main use cases for AI in manufacturing?
The established use cases for AI in manufacturing are predictive maintenance, visual quality inspection, process optimization, demand and supply planning, and generative design. The newer one, made possible by multimodal models that can reason about video, is real-time work guidance: a worker wears a camera, the model sees what they see, and it coaches them through the job step by step. Y Combinator's example is a model telling a worker to turn off a valve, use a specific wrench, and replace a part that looks worn. Guidance is the use case with the shortest path to a live pilot, because it requires no data from the machine and no change to the equipment.
Is AI replacing manufacturing workers?
AI is not replacing manufacturing workers at the pace the headlines suggest, and the constraint is capability rather than willingness. AI systems still cannot reliably act in an unstructured physical environment, which is why the deployments landing today guide a human rather than replace one. The bigger force in manufacturing is the opposite of replacement: Deloitte and The Manufacturing Institute project that US manufacturing could need as many as 3.8 million new employees between 2024 and 2033, with roughly 1.9 million positions at risk of going unfilled. The scarce input is not headcount alone, it is the embedded judgment of experienced operators, and guidance systems exist mainly to deliver that judgment to people who have not yet acquired it.
How do you start an AI project in a factory?
Start with one job on one line rather than an industry, and pick the layer where being wrong is cheap. A wrong recommendation costs a conversation with a supervisor, while a wrong actuation costs a part, a line, or a person, so an advisory pilot clears a safety review that a robotic cell would not. Before writing code, get three things from the design partner. Get the network diagram and the operational technology (OT) security contact, because plant floors are segmented under standards like ISA/IEC 62443, and direct cloud calls from the control layer are a design decision rather than an oversight. Get the list of people who can veto the project, which usually includes environmental health and safety (EHS), the OT security team, quality, and in many facilities a union or works council. Get the specific metric your product moves and the name of the person who owns it.
Build your AI Operating System
A practical course to grow with AI, build internal tools, and operate safely. Join the waitlist and you'll be first in when the course opens.