Reliability intelligence
AI incident investigation for autonomous systems.
Pharion reconstructs failures across ROS, telemetry, sensors, and robot state - identifying what happened, why it happened, and how to prevent it.
Built for robotics engineering teams operating autonomous fleets.
Investigation of incident 042: LiDAR degradation leads to localization instability, planner hold, and a protective stop on AMR-17.
- .04
- .21
- .47
- .63
- .71
- .89
Localization instability caused by degraded LiDAR observations. Evidence traces from return density through pose covariance to the protective stop.
The problem
When autonomous systems fail, engineers become detectives.
A single incident can span bags, telemetry, cameras, LiDAR, IMU, localization, planner state, controller output, and logs. Reconstructing causality still happens by hand.
ROS / MCAP
bags · topics · tf
Telemetry
rates · estop · power
Cameras
streams · frames
LiDAR
returns · intensity
IMU
accel · gyro
Localization
pose · covariance
Planner
costmap · trajectory
Controller
cmd_vel · safety
System logs
nodes · diagnostics
An engineer opens each source in a different tool, aligns clocks by hand, and guesses which events belong to the same failure.
Reached hours later - if the trail can be reconstructed at all.
The data exists. The intelligence to connect it doesn't.
Product
Observability tells you what happened.
Pharion investigates why.
Visualization and monitoring remain essential. Engineers still have to stitch timelines, compare robot state, and argue about causality. Pharion sits above that stack - it does not replace it.
- 01InfrastructureRobots, fleets, and the ROS / ROS 2 runtime already producing state.
- 02DataMCAP, bags, telemetry, LiDAR, cameras, IMU, GPS, logs, diagnostics.
- 03ObservabilityFoxglove, RViz, Grafana, and logs show what the system did. Necessary. Not an investigation.
Timeline
Reconstructed event clock
Evidence
Observable system behavior
Root cause
Why the failure occurred
Recommendation
What to change next
How it works
Ingest. Reconstruct. Investigate. Explain.
Connect the streams the robot already emits, align them to one clock, find the causal chain, and ship a recommendation the team can act on.
01Ingest
Connect the streams the robot already emits: ROS / ROS 2, MCAP, telemetry, cameras, LiDAR, IMU, GPS, diagnostics, logs.
- MCAP / bagsamr17_1432.mcap
- LiDAR/scan · intensity
- Localizationamcl · pose cov.
- Plannercostmap · traj
- Controllercmd_vel · estop
- Logsnodes · diagnostics
02Reconstruct
Align sensors, localization, planner state, controller output, and software events onto one incident clock.
- 11.04LiDAR
- 11.21Loc
- 11.47Pose
- 11.71Safety
- 11.89Stop
03Investigate
Detect anomalies and ordered causal links - not co-plotted charts, but which event caused the next.
LiDAR density ↓
loc confidence
Pose covariance ↑
planner hold
Unstable pose
safety stop
04Explain
Emit an evidence-backed root cause and the corrective action the team can ship.
LiDAR degradation → unstable pose → protective stop.
Raise return-density watchdog; hold planner on loc quality < 0.6.
Investigation
An investigation, not a dashboard.
Incident #042 - AMR-17 stopped in aisle B4. Timeline, robot state, evidence, and recommendation live in one workspace, each finding pointing at observable behavior.
Workspace for incident 042: LiDAR degradation leads to localization instability, planner hold, and a protective stop. Focus or select a subsystem, chain step, or evidence item to see related events.
Localization instability caused by degraded LiDAR observations.
Raise the LiDAR return-density watchdog on AMR-17. Gate planner execution when localization quality drops below 0.6. Replay against the last known-good map at 14:31:40.
A demo walks the same reconstruction on your fleet data.
Evidence
Don't just get an answer. Get the evidence.
Every conclusion traces to observable system behavior. If it cannot be shown in the data, it does not ship as a finding.
Localization instability caused by degraded LiDAR observations.
Incident timeline
The story of a failure, reconstructed.
Independent subsystems, one clock. From 14:32:11.04 to 11.89 the sequence is visible - not inferred from memory.
- .04
- .21
- .47
- .63
- .71
- .89
Where it applies
Autonomous systems where failure investigation matters.
The same investigation problem exists anywhere a machine acts without a human in the loop. We are not locked to a first vertical.
Warehouse AMRs
Unexplained stops in aisles, docks, and congestion.
Autonomous forklifts
Load, lift, and path failures that halt a shift.
Manufacturing cells
Motion, perception, and safety events on the line.
Drones & defense
Aerial systems where after-action reconstruction cannot wait on tribal knowledge.
Space robotics
Remote platforms where replay is expensive and nobody can walk up to the machine.
Advanced robotics
Humanoids, field, inspection, and agricultural systems - dense state, high cost of an unexplained fault.
Why now
The more autonomous systems scale, the more expensive unexplained failures become.
- 01Systems are entering the real worldRobots leave prototypes and enter live operations - warehouses, yards, fields, and airspace - where failures have a cost.
- 02Sensors, software, and models share the vehiclePerception, planning, control, and learned components fail across subsystem boundaries. The trail is no longer in one log.
- 03Reliability is the advantageAs fleets grow, unexplained downtime compounds. Teams that reconstruct causality spend less time guessing.
Vision
The reliability intelligence layer for autonomous systems.
- Engineers investigate failures by hand - bags, logs, telemetry, and tribal knowledge.
- Autonomous systems continuously understand their own failures, with evidence attached.
- Every robot has a memory of what happened, why it happened, and what changed.
Team
Built across robotics, autonomy, and infrastructure.
Experience across robots in the field, autonomy stacks, and the software that keeps fleets running. About us · Contact us.
- Robotics
- Autonomy
- AI systems
- Software systems
- Engineering infrastructure
When your autonomous system fails, know why.
Fragmented data becomes an evidence-backed investigation - what happened, why, and what to change.
- Logs
- Telemetry
- Sensors
- Robot state
- Investigation
