Do Not Certify the Machine. Certify the Evidence.
The first three parts lead to a larger conclusion.
If AI is still developing, and if current systems may represent an early phase of a trajectory toward more general and potentially superhuman intelligence, then the safety infrastructure must mature before those systems become deeply embedded in physical environments.
We should not wait until AI reaches a much higher level of capability to ask how it will coexist safely with humanity.
We should build the measurement infrastructure now.
The goal is not to declare that every machine is either “safe” or “unsafe.” Safety is not a permanent label. It is a claim that must be tied to a specific system, a defined operating environment, particular tasks, known failure modes, and a continuing body of evidence.
That is why the central principle should be:
Do not certify the machine. Certify the evidence.
Standardize Evidence, Not Innovation
Manufacturers should not necessarily be required to use the same:
Algorithms
Models
Sensors
Training datasets
Hardware architectures
Development methodologies
Control systems
Innovation depends on allowing multiple technical approaches.
But the evidence required before deployment can be more standardized.
A public-facing embodied AI system could be required to demonstrate safety through a layered evidence process.
The final certification should not claim:
“This robot is safe.”
That statement is too broad to be scientifically meaningful.
A stronger and more defensible conclusion would be:
This system demonstrated specified safety properties across defined populations, environments, behaviors, failure modes, and consequence classes within a defined operating envelope.
That language acknowledges the reality of engineering: performance depends on conditions.
The Human State Space Problem
The world contains billions of people. It may be tempting to frame the challenge as a question of whether a system needs to identify every person before operating safely in public.
But the deeper issue is not merely the number of individual identities.
It is the size of the human state space.
One individual may:
Walk, run, crouch, dance, fall, or stop suddenly
Carry an object or push a stroller
Use a wheelchair, cane, walker, guide dog, or prosthetic
Change direction without warning
Behave differently when distracted, tired, frightened, or rushed
Interact with other people in groups
Speak, gesture, signal, or communicate in unexpected ways
Encounter the system under unusual lighting, weather, crowding, or noise conditions
Now multiply that by billions of people, countless environments, different cultures, varying physical characteristics, unexpected events, and system failures.
The possible combinations become effectively unbounded.
That means the practical goal cannot be:
Test every human.
The goal must be:
Demonstrate that the system generalizes safely beyond the people, behaviors, environments, and circumstances used to develop it.
What a Serious Safety Case Requires
A credible safety case for public embodied AI should include more than a successful demo or a favorable internal test report.
It should include:
Representative population testing across meaningful human variation
Independent evaluation rather than only manufacturer self-assessment
Blind testing where evaluators do not know the system’s expected response
Scenario generation for rare, ambiguous, and high-consequence events
Testing for intentional misuse and accidental disruption
Simulated and real-world failure scenarios
Clear limits on the system’s approved operating environment
Human oversight requirements proportionate to uncertainty and consequence
Incident reporting, near-miss reporting, and post-deployment monitoring
A mechanism for reducing or suspending authority when evidence no longer supports safe operation
The most important question is not:
Have we tested enough people?
It is:
What evidence demonstrates that the system remains safe when it encounters people it has never seen before?
Intelligence, Restraint, and Adaptability
Everything in this series returns to one central question:
Can the system safely interact with humans who have not been selected, trained, modeled, or instructed to interact with it?
If the answer is not supported by evidence, the solution is not necessarily to abandon the technology.
It is to reduce the deployment envelope.
Industrial robotics can continue. Commercial experimentation can continue. Supervised public pilots can continue. Testing can continue. Learning can continue.
But uncontrolled public deployment should meet a distinct evidentiary threshold.
As machine intelligence becomes more capable, physical systems will increasingly be able to act upon the world.
That creates a fundamental distinction:
Intelligence is capability.
Safety is restraint.
Human integration is adaptability.
Superintelligence—if and when it emerges—will require all three.
The objective is not to make machines perfectly understand every human. That may be impossible.
The objective is to ensure that imperfect understanding does not produce unacceptable consequences.
The future should not require humanity to learn how to behave like a machine.
It should require machines to become capable of safely existing within the extraordinary variability of humanity.
The public should not be the test track.
And we should not wait for superintelligence to arrive before building the infrastructure required to live alongside it.
We are the people and the place.

