Writings, Regrets, and Re-skillings during the AI Revolution

Tag: ite135

  • Four Rules for Governing World Models

    Series note. To mark the completion of my ECE 5864 Critical Engineering capstone, I’m publishing a seven-part adaptation for a special week of ITE 135: AI Awareness. These posts are longer and more formal than my usual entries because they began as a graduate research paper. This is Part 7 of 7.


    Critical Engineering Recommendations

    Require independent safety and quality testing before deployment. A world-model robot controller should not be used around workers until an independent evaluator has tested the exact model, robot, task, and factory environment. Testing should include unusual parts, blocked sensors, lighting changes, worker entry into the cell, damaged equipment, and other foreseeable failures. A successful vendor demonstration geared toward marketing and the venture capitalists is not enough. The system should also be retested after major model updates, new training data, or changes to the environment. If the system cannot show safe behavior under these conditions, it should not be used in production environments.

    Protect workers, jobs, and the right to stop the system. Employers should complete a risk and impact assessment before introducing this type of automation. The assessment should estimate jobs lost, jobs changed, wages affected, training needs, and new safety hazards. Workers and their representatives should participate before deployment, not after the purchasing decision has already been made. Factories should have a clearly tested emergency-stop process, and workers should have the authority to halt operation without retaliation. Employers should also fund retraining, continued manual practice, and transition support for workers whose jobs are eliminated or reduced. Productivity gains should not be treated as a complete benefit calculation while job loss and lost wages are externalized by the company.

    Do not count a human monitor as a safety control. The most common industry answer to the risks described above is to place a trained person in a supervisory position and treat that person as the safeguard. The Avride investigation is direct evidence that this does not work. A qualified safety operator occupied the driver’s seat during all sixteen crashes under review, and in only one of them did that operator attempt to intervene. A learned model/agent acts faster than a passive observer can reliably react, and a supervisor watching a system that behaves correctly most of the time will stop watching it closely. Employers and regulators should therefore require that a deployment be safe on its engineered limits alone, e.g., force and speed ceilings, protective stops, physical separation, and restricted task boundaries that hold whether or not anyone intervenes. Where human intervention is still claimed as a mitigation, it should be measured rather than assumed: the assessment must document the time a worker actually needs to detect a fault and stop the machine, verified under realistic conditions rather than during a scheduled test. Supervision should be treated as a way to catch problems and improve the system over time, never as the primary barrier between a machine-learning model and a human body.

    Assign responsibility and blame before an accident. Contracts and regulations should identify which party is responsible for the foundation model, robot integration, safety controls, software updates, and day-to-day operation. Each deployed version should have a model data provenance card that records its origin, training changes, known limits, and safety tests. Serious failures and near misses should be reported to a shared public database so that one company’s accident can prevent another company’s possible fatality. Regulators should also require secure update procedures and checks for compromised model weights because research has shown that generative pipelines can be poisoned through hidden backdoors.

    World models may eventually improve reliability, but that possibility does not justify testing their emerging capabilities on workers in the field as the WEF suggests. Our policy goal should not be to stop world-model research. It should be to prevent companies from shifting the risks of experimental automation onto people who did not design the system and may not share in its profits. Independent benchmarks and quality testing, enforceable worker protections, job-impact planning, public incident reporting, and clear liability laws would require the evidence to catch up with the runaway marketing and hype. Until tech companies and their capitalist advocates can demonstrate that these systems are safe in the real-world physical environments where they will operate, they should not be permitted to gamble with workers’ jobs, property, or lives.

  • The Policy Problem of Physical AI

    Series note. To mark the completion of my ECE 5864 Critical Engineering capstone, I’m publishing a seven-part adaptation for a special week of ITE 135: AI Awareness. These posts are longer and more formal than my usual entries because they began as a graduate research paper. This is Part 6 of 7.


    Introduction and Context

    World models are AI systems that try to predict how the physical world will change. In a factory, that prediction can become a robot action. This makes the technology more dangerous than an ordinary generative AI system. A bad chatbot answer may ‘hallucinate’ or confuse someone. A bad robot command may crush a hand, damage equipment, or kill someone. FLUX-mimic is already being tested at Audi, even though its reliability has not been independently established. The main policy problem is therefore simple: companies are moving experimental AI from video generation into industrial control before workers, regulators, and the public have reliable ways to judge whether it is safe.

    Relevance to Policy and Practice

    Policy action is needed because the costs and benefits will fall on different groups. Audi and its vendors may gain faster production, lower labor costs, and a new market for foundation models. Workers may face injury, job loss, lower wages, or reassignment to monitoring roles with less skill and bargaining power. The safety risk is already real and urgent. NIOSH identified 41 robot-related deaths in the United States between 1992 and 2017, and it warns that unexpected contact, crushing, trapping, and concern about job loss remain important risks as robots work closer to people (CDC 2024). In May 2026, federal regulators opened a defect investigation into Avride, an Uber robotaxi partner operating in Austin and Dallas, after sixteen crashes in which the automated driving system merged into occupied lanes, failed to slow for traffic ahead, and struck objects partly blocking the road (O’Kane 2026). Interestingly, a trained human monitor sat in the driver’s seat for every one of them, and in only a single incident did that monitor attempt to take over. The regulator’s warning about “inappropriate assertiveness and insufficient competence” provokes the very question this report is concerned with: can a world model remain safe at the moment an ordinary reality stops resembling its training data?

    Current world models add another problem: their actions may be difficult to predict and control because the models learn stochastic patterns instead of following a complete set of explicit rules. Symbolic guardrails and explicit prohibitions at some point become necessary again to avoid worst-case statistical blunders. Public stakeholders should have the right to participate in the deliberation and articulation of such absolute prohibitions. These protections remain necessary, but a learned controller raises questions that ordinary guarding and emergency stops cannot answer. Regulators also need to know how the model was trained, how often it fails, what changes after fine-tuning, and whether it behaves safely outside its demonstration data. The NIST AI Risk Management Framework can help organizations identify and manage these risks, but it is voluntary and does not itself prevent an unsafe system from entering a factory (NIST 2023). In short, the current system relies too heavily on vendors and employers to judge their own technology.

    References

    CDC (Centers for Disease Control and Prevention). 2024. “Robotics in the Workplace: An Overview.” National Institute for Occupational Safety and Health, February 9, 2024.

    NIST (National Institute of Standards and Technology). 2023. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. Gaithersburg, MD: National Institute of Standards and Technology.

    O’Kane, Sean. 2026. “Uber Partner Avride Is under Investigation for Self-Driving Crashes.” TechCrunch, May 8, 2026.

  • When a World Model Makes the Wrong Physical Decision

    Series note. To mark the completion of my ECE 5864 Critical Engineering capstone, I’m publishing a seven-part adaptation for a special week of ITE 135: AI Awareness. These posts are longer and more formal than my usual entries because they began as a graduate research paper. This is Part 5 of 7.


    Risks of World Models

    Risks of deploying world models at scale in physical robotics include possible death, injury, and damage. A robot controlled by a world model can make the wrong physical decision, injure or kill a worker, damage equipment, or stop a production line. The same system may also eliminate jobs or reduce workers to monitoring a process they no longer fully understand. These risks are harder to manage because FLUX-mimic is still a vendor-led production test, not an independently validated industrial control system. The critical question is therefore not whether the model can produce convincing demonstrations (it can). It is whether the system is safe, reliable, and accountable when something unexpected happens around real people and machinery.

    The most immediate engineering risk is unreliable behavior. A world model can produce an answer that looks reasonable while being physically wrong (WEF 2026). The literature survey in the Technical Summary section already showed that current video models can generate realistic scenes that fail the fundamental laws of gravity, fluids, and heat (Ding et al. 2025). That problem becomes far more serious when the output is a robot command instead of a video frame. A bad prediction could cause a robot to grip too hard, move into a worker’s space, drop a component, or damage a vehicle. Industrial robots have caused fatal accidents before, especially during maintenance, testing, setup, and other situations when workers enter the robot’s operating area (CDC 2024).

    Another possible risk is job loss. The main economic purpose of this system is to automate assembly work that previously required people. If it succeeds, some repetitive and injury-prone jobs may disappear, but so may stable manufacturing jobs that support workers and communities. The benefits will not be shared automatically. Audi and its technology vendors may gain productivity and lower labor costs while workers face layoffs, reassignment, or pressure to accept lower-skill monitoring roles. Retraining may help some workers, but it is not a complete answer if there are fewer comparable jobs available. Any serious evaluation should therefore count displaced workers and lost wages as costs, not treat them as side effects outside the engineering problem.

    Automation can also weaken human oversight over time. Experienced workers often notice small changes in sound, motion, resistance, or part quality before a formal alarm appears. If people stop performing the task, they may gradually lose the practical knowledge needed to recognize when the robot is behaving dangerously. Supervisors could then become passive monitors of a system they cannot fully inspect or explain. For human oversight to mean anything, workers must retain the authority to stop the robot, receive training on its known failure modes, and regularly practice the manual skills needed to take over safely.

    Finally, the number of supply-chain partners can create a diffusion of irresponsibility effect when a serious failure occurs. Black Forest Labs trained the foundation model FLUX 3, mimic adapts it for robot control via FLUX-mimic, Audi places it on the factory floor, and workers are exposed to the result. If the robot injures someone, each organization may point to another part of the supply chain. Open-weight releases may make this problem worse because many companies can create modified versions with different data, safeguards, and quality controls. Research has also shown that generative pipelines can be deliberately compromised through backdoors and data poisoning (Lapid and Dubin 2025). Clear responsibility, documented model changes, independent safety testing, and incident reporting should therefore be minimum conditions for deployment—not add-ons added after a fatal accident occurs.

    References

    CDC (Centers for Disease Control and Prevention). 2024. “Robotics in the Workplace: An Overview.” National Institute for Occupational Safety and Health, February 9, 2024.

    Ding, Jingtao, Yunke Zhang, Yu Shang, Yuheng Zhang, Zefang Zong, Jie Feng, Yuan Yuan, et al. 2025. “Understanding World or Predicting Future? A Comprehensive Survey of World Models.” ACM Computing Surveys 58 (3): Article 57.

    Lapid, Raz, and Almog Dubin. 2025. “Backdoors in Conditional Diffusion: Threats to Responsible Synthetic Data Pipelines.” AAAI 2026 Workshop on Shaping Responsible Synthetic Data in the Era of Foundation Models.

    WEF (World Economic Forum). 2026. Top 10 Emerging Technologies of 2026. Geneva: World Economic Forum, pp. 29–31 and 47–48.

  • A World Model Enters the Factory

    Series note. To mark the completion of my ECE 5864 Critical Engineering capstone, I’m publishing a seven-part adaptation for a special week of ITE 135: AI Awareness. These posts are longer and more formal than my usual entries because they began as a graduate research paper. This is Part 4 of 7.


    Use Case Description

    For a recent use case analysis, timely as of summer 2026, I chose FLUX-mimic, a new world-model robotic system undergoing production testing at Audi factories for the installation of flexible door seals and the handling of car components in final vehicle assembly (Black Forest Labs 2026b). The system is useful for this report because its learned video-action policy is derived from a multimodal foundation model trained on large-scale video, vision, language, and robotics data rather than being specified entirely through task-specific rules (Black Forest Labs 2026a). That learned policy is only one component of a complete robot-control system that also includes sensors, middleware, action decoding, actuators, conventional software, and safety controls. The relevant engineering question is therefore not whether the model acts alone, but whether the entire robot cell remains safe when the learned component makes an incorrect prediction.

    Black Forest Labs’ wager in its public announcement “Real World Models” is that a generative-AI world model trained to represent reality convincingly must implicitly learn some of reality’s physical regularities, such as how a cable droops or falls, and that this video-learned knowledge can be redirected through transfer learning from pixel generation to motor control. That places the system squarely in the category that the World Economic Forum’s 2026 report describes as world models, alongside the NVIDIA Cosmos platform it highlights (WEF 2026). Adopting Ding et al.’s taxonomy, FLUX-mimic is a future-prediction world model derived partly from generative video AI.

    The main use actors are as follows: Audi’s manufacturing engineers select candidate workflows that have not previously been automated, such as tasks involving floppy or deformable parts, in which geometry is not fixed and strictly repeated motion usually fails. Robotics engineers then fine-tune the pretrained FLUX 3 foundation model into a “video-action model” for each selected task using human demonstrations. The vendors describe this adaptation as requiring only a limited amount of task-specific demonstration data rather than extensive conventional programming (Black Forest Labs 2026b). The adapted model is deployed to a robot cell, where it consumes camera streams and outputs control actions at operational speed. Audi’s process engineers then deploy the cell on the production assembly line, while line technicians supervise operation and handle exceptions. Behind this sequence stands an under-discussed fourth actor: the pretraining pipeline and the broader media economy that supplied its data. Black Forest Labs describes training at massive scale but does not fully disclose the provenance, licensing status, or consent conditions governing that material. This opacity creates unresolved copyright and labor concerns, particularly because the same company sells video generation into film and advertising markets that may disrupt work traditionally performed by human crews and performers.

    Benefits

    The clearest near-term benefit is the automation of physical tasks that conventional industrial robots handle poorly. Traditional systems excel when parts are rigid, locations are fixed, and motions can be repeated exactly; they are far less capable with seals, cables, fabric, food, and other deformable or fragile materials whose shape changes during handling. A world-model controller that predicts those changes could extend automation into final-assembly work that remains manual, repetitive, and ergonomically difficult. In Audi’s use case, successful seal installation could therefore produce both a productivity benefit and a worker-safety benefit by reducing sustained force, awkward posture, and repetitive motion provided that the new robotic system does not introduce greater hazards of its own. A related benefit is faster adaptation and safer experimentation. Conventional automation may require new fixtures, task-specific programming, and extensive validation whenever a product or workflow changes. FLUX-mimic’s vendors claim that its controller can instead be adapted quite quickly compared to traditional robotics (re)programming.

    The possible impact also extends well beyond manufacturing. In transportation, world models could help vehicles anticipate how traffic participants and road conditions may evolve, compare possible maneuvers, and rehearse uncommon emergencies before deployment. In medicine, related predictive models could support treatment planning, medical imaging, surgical rehearsal, or assistive robotics by estimating how a patient or procedure may respond under different interventions. In hazardous-materials handling and disaster response, robots could explore contaminated, explosive, radioactive, or structurally unstable environments while keeping people farther from immediate danger. World models have also been proposed for climate and weather scenarios, scientific modeling, and other systems in which direct experimentation is costly or impossible (Ding et al. 2025). Across these fields, the common benefit is not autonomous “understanding” in the strong sense, but the ability to compare possible futures before choosing an action in the real world.

    References

    Black Forest Labs. 2026a. “FLUX 3 — Real World Models.” Black Forest Labs, July 23, 2026.

    Black Forest Labs. 2026b. “FLUX 3 x mimic: The Next Generation of Video-Action Models.” Black Forest Labs, July 23, 2026.

  • Plausible Physics Is Not Physical Understanding

    Series note. To mark the completion of my ECE 5864 Critical Engineering capstone, I’m publishing a seven-part adaptation for a special week of ITE 135: AI Awareness. These posts are longer and more formal than my usual entries because they began as a graduate research paper. This is Part 3 of 7.


    Assessment and Analysis

    Read critically, Ding et al’s most consequential contribution is the evidence it compiles against the field’s own marketing and hype. The product category is named for understanding the world, yet the authors document generative AI systems failing precisely at physical reasoning while succeeding at visual plausibility. That gap is diagnostic of what these systems optimize: a video generator is trained to make the next frame look right to a human eye, so gravity, friction, and thermal behavior are learned only insofar as they leave visible traces in pixels. Internal coherence is rewarded; correspondence to the world is, at best, a by-product. This gap could produce a new category of physical AI failure: not merely videos with broken physics, but robots whose plausible-looking predictions generate unsafe behavior. A hallucination on a factory floor, not to mention in a school, hospital, or road environment, may cost lives.

    Their own evidence also licenses a critique of the survey’s own central move. Ding and colleagues resolve the field’s definitional dispute by declaring understanding and prediction to be two functions of a single underlying idea. It is an elegant synthesis, but I think it is too slippery still. Current laboratory studies and frontier demonstrations suggest that the two functions come apart in practice: a system can be excellent at generating plausible futures while holding no rule-governed model of the present. If that is so, then ‘understanding’ and ‘predicting’ are not two aspects of one capability but two distinct claims. The shared slippery label permits a vendor to demonstrate the second while marketing the first.

    The survey’s “missing benchmarks” observation compounds this problem, and for the developer audience it is arguably the paper’s most important point: deployment of world-model systems is proceeding without an agreed method for validating opaque learned controllers across distribution shifts, model updates, and changing operating conditions. Industrial robotics is not devoid of safety governance; existing machinery standards address robot cells, guarding, protective functions, integration, and foreseeable human access. The unresolved gap is narrower but consequential: these instruments do not yet provide a canonical, model-specific method for establishing whether a learned controller remains reliable after fine-tuning, software updates, sensor drift, or exposure to conditions outside its training data. Without such evidence, a vendor demonstration can become a de facto performance test whose task, conditions, and footage are selected by the vendor.

    Two further limitations of the survey should be named. As a survey, the paper produces no original experiment; its findings are curated from others’ work, and its authority rests on the quality of its taxonomic synthesis rather than on evidence it generated itself. In a field moving this quickly, any survey’s shelf life is also short, and this one already predates several systems now central to public discussion, including the FLUX 3 lineage I take up in the next installment, my Use Case Analysis.

    Ding and colleagues provide definitional clarity and the taxonomy that discussion of world models has so far lacked. The state of the field they document nevertheless remains premature: researchers have demonstrated potential, but not safety or reliability, and quality-oriented benchmarks remain fragmented. Once again, Silicon Valley’s “move fast and break things” ethos is entering a high-impact domain with physical consequences before civil society and government institutions have developed adequate model-specific assurance and accountability mechanisms.

    Reference

    Ding, Jingtao, Yunke Zhang, Yu Shang, Yuheng Zhang, Zefang Zong, Jie Feng, Yuan Yuan, et al. 2025. “Understanding World or Predicting Future? A Comprehensive Survey of World Models.” ACM Computing Surveys 58 (3): Article 57.

  • What Is a World Model?

    Series note. To mark the completion of my ECE 5864 Critical Engineering capstone, I’m publishing a seven-part adaptation for a special week of ITE 135: AI Awareness. These posts are longer and more formal than my usual entries because they began as a graduate research paper. This is Part 2 of 7.


    Technical Summary

    A. Introduction

    In “Understanding World or Predicting Future? A Comprehensive Survey of World Models,” Jingtao Ding and eleven colleagues at Tsinghua University (published in ACM Computing Surveys 58, 3, 2025) provide the first comprehensive peer-reviewed survey of world-model research. Its central move is to resolve the term’s contested definition by organizing the entire literature around two functions: i. building internal representations that help a machine understand the world and ii. predicting future states of the world to guide decisions.

    B. Description

    The problem that Ding et al. address is that “world model” has come to mean different things to different research communities, all hiding beneath the marketing hype. One lineage, descending from Ha and Schmidhuber’s 2018 work and from model-based reinforcement learning, treats a world model as a compressed internal representation that an agent consults to make decisions. Ha and Schmidhuber used learned latent representations to produce the first reported solution of the CarRacing benchmark. In a separate VizDoom experiment, they trained a policy entirely inside an environment generated by the model and transferred that policy successfully to the original game (Ha and Schmidhuber 2018).

    A second lineage, energized by video generators such as Sora and FLUX, treats a world model as a simulator that renders plausible future states of the physical world. The “Joint Embedding Predictive Architecture,” a.k.a. JEPA, proposed by Yann LeCun, former Chief AI Research Scientist at Meta, argues that a machine should both represent the world abstractly and use that representation to imagine outcomes before acting. The embedding in that name points to a mechanism both lineages increasingly share: instead of learning only from text data or reproducing every image pixel, many current systems compress video, lidar, motion, and other sensor data into a ‘latent’ or abstract space, which consists of numerical vector representations intended to preserve important event information and dynamic relationships while omitting overwhelming raw detail. Prediction then happens in that compressed space rather than at the pixel level, which is what makes the approach computationally tractable at scale, but also what makes its omissions consequential.

    The authors conduct a systematic review in three passes. First, they survey the internal-representation branch: world models inside model-based reinforcement learning agents, and the more recent finding that large language models (LLMs) appear to acquire latent world knowledge, including spatial and temporal structure, without being explicitly trained for it. Second, they survey the future-prediction branch, tracing a progression from video generation toward interactive embodied environments, from producing footage of the world to producing worlds an agent can act in. Third, they examine three application domains (autonomous driving, robotics, and social simulacra), showing that each domain draws on their two functions in different proportions. The authors attempt to unify these two fields as two functions of one underlying idea: understanding the present and predicting the future.

    The paper closes with an open problem and findings that are largely negative. Current video generators achieve visual realism but fail systematic tests of physical law: the authors cite diagnostic studies showing errors in gravity, fluid, and thermal dynamics, and evidence that scaling produces case-based rather than rule-based generalization. So far world models ‘memorize’ situations without inducing the physics beneath them. Finally, the authors document that evaluation is fragmented: because the field pursues divergent goals with heterogeneous methods, no canonical benchmark or metric for “being a good world model” yet exists.

    References

    Ding, Jingtao, Yunke Zhang, Yu Shang, Yuheng Zhang, Zefang Zong, Jie Feng, Yuan Yuan, et al. 2025. “Understanding World or Predicting Future? A Comprehensive Survey of World Models.” ACM Computing Surveys 58 (3): Article 57.

    Ha, David, and Jürgen Schmidhuber. 2018. “Recurrent World Models Facilitate Policy Evolution.” In Advances in Neural Information Processing Systems 31 (NeurIPS 2018), 2450–2462.

  • World Models: The Next Frontier of AI Hype

    Series note. To mark the completion of my ECE 5864 Critical Engineering capstone, I’m publishing a seven-part adaptation for a special week of ITE 135: AI Awareness. These posts are longer and more formal than my usual entries because they began as a graduate research paper. This is Part 1 of 7.


    Introduction

    World models are AI models that form—through massive and deep machine learning on video, vision, language, and robotics data—an internal representation of a physical environment’s state and dynamics to predict how that environment may change and evaluate possible actions within it. They are a key enabling approach for an emerging generation of physical-AI systems and robotics. They are also the latest frontier of Silicon-Valley-inspired hype and market speculation because they could enable so-called ‘agentic AI’ to control autonomous machines in real-world environments by making inferences, selecting actions, and predicting their effects. If the technology proves viable at scale, world models could enable robots, cobots, autonomous vehicles, and other physical-AI systems to replace some human labor or better collaborate with human workers. The potential impact is immense but remains largely speculative at present.

    Last month, the World Economic Forum selected world models as one of its Top 10 Emerging Technologies of 2026. Its report presents them as a foundation for AI systems that can plan, reason about, and act autonomously in physical settings, including robotics, industrial operations, climate modeling, and transportation (WEF 2026, 29–31). Unfortunately, their business-facing report glosses over a central critical-engineering risk: a world model may be coherent within its own representation yet still grossly misunderstand the world, wreaking havoc and harm in the environment. Critical errors and ‘bug’ defects may remain hidden until a world model is actually deployed and encounters real people/places/things outside its training environment.

    World models are an active area of research, development, and technology investment. Major engineering obstacles remain, including model degradation over time, stringent reliability requirements for systems operating around people and property, computational efficiency at scale, and high energy and environmental costs. Terminological confusion and a proliferation of hype-driven buzzwords further complicate analysis and assessment. None of these uncertainties has dampened enthusiasm: Silicon Valley is wagering heavily and publicly that world models will become both deployable and lucrative in the near future.

    This critical-engineering report is written for developers, technology managers, workers, and other public stakeholders who may be affected by or curious about emerging world-model and physical-AI systems. In the next installment, I summarize the most comprehensive peer-reviewed survey to date of the field’s competing definitions, main technical approaches, major application domains, and unresolved problems. I argue that “world model” currently functions as both a technical category and a marketing device. Vendors can slide rhetorically from a model’s ability to generate plausible representations to claims that it understands physical reality well enough to act safely in settings such as schools, malls, and elder care. Although world models show substantial promise for representing, predicting, and controlling spatiotemporal environments, limited demonstrations do not establish the calibrated reliability and accountability required in such settings. Given what is at stake in the natural and social world, these systems demand critical scrutiny before broad deployment.

    References

    WEF (World Economic Forum). 2026. Top 10 Emerging Technologies of 2026. Geneva: World Economic Forum, pp. 29–31 and 47–48.