Series note. To mark the completion of my ECE 5864 Critical Engineering capstone, I’m publishing a seven-part adaptation for a special week of ITE 135: AI Awareness. These posts are longer and more formal than my usual entries because they began as a graduate research paper. This is Part 7 of 7.
Critical Engineering Recommendations
Require independent safety and quality testing before deployment. A world-model robot controller should not be used around workers until an independent evaluator has tested the exact model, robot, task, and factory environment. Testing should include unusual parts, blocked sensors, lighting changes, worker entry into the cell, damaged equipment, and other foreseeable failures. A successful vendor demonstration geared toward marketing and the venture capitalists is not enough. The system should also be retested after major model updates, new training data, or changes to the environment. If the system cannot show safe behavior under these conditions, it should not be used in production environments.
Protect workers, jobs, and the right to stop the system. Employers should complete a risk and impact assessment before introducing this type of automation. The assessment should estimate jobs lost, jobs changed, wages affected, training needs, and new safety hazards. Workers and their representatives should participate before deployment, not after the purchasing decision has already been made. Factories should have a clearly tested emergency-stop process, and workers should have the authority to halt operation without retaliation. Employers should also fund retraining, continued manual practice, and transition support for workers whose jobs are eliminated or reduced. Productivity gains should not be treated as a complete benefit calculation while job loss and lost wages are externalized by the company.
Do not count a human monitor as a safety control. The most common industry answer to the risks described above is to place a trained person in a supervisory position and treat that person as the safeguard. The Avride investigation is direct evidence that this does not work. A qualified safety operator occupied the driver’s seat during all sixteen crashes under review, and in only one of them did that operator attempt to intervene. A learned model/agent acts faster than a passive observer can reliably react, and a supervisor watching a system that behaves correctly most of the time will stop watching it closely. Employers and regulators should therefore require that a deployment be safe on its engineered limits alone, e.g., force and speed ceilings, protective stops, physical separation, and restricted task boundaries that hold whether or not anyone intervenes. Where human intervention is still claimed as a mitigation, it should be measured rather than assumed: the assessment must document the time a worker actually needs to detect a fault and stop the machine, verified under realistic conditions rather than during a scheduled test. Supervision should be treated as a way to catch problems and improve the system over time, never as the primary barrier between a machine-learning model and a human body.
Assign responsibility and blame before an accident. Contracts and regulations should identify which party is responsible for the foundation model, robot integration, safety controls, software updates, and day-to-day operation. Each deployed version should have a model data provenance card that records its origin, training changes, known limits, and safety tests. Serious failures and near misses should be reported to a shared public database so that one company’s accident can prevent another company’s possible fatality. Regulators should also require secure update procedures and checks for compromised model weights because research has shown that generative pipelines can be poisoned through hidden backdoors.
World models may eventually improve reliability, but that possibility does not justify testing their emerging capabilities on workers in the field as the WEF suggests. Our policy goal should not be to stop world-model research. It should be to prevent companies from shifting the risks of experimental automation onto people who did not design the system and may not share in its profits. Independent benchmarks and quality testing, enforceable worker protections, job-impact planning, public incident reporting, and clear liability laws would require the evidence to catch up with the runaway marketing and hype. Until tech companies and their capitalist advocates can demonstrate that these systems are safe in the real-world physical environments where they will operate, they should not be permitted to gamble with workers’ jobs, property, or lives.