Skip to content
PHARION

Reliability intelligence

AI incident investigation for autonomous systems.

Pharion reconstructs failures across ROS, telemetry, sensors, and robot state — identifying what happened, why it happened, and how to prevent it.

Built for robotics engineering teams operating autonomous fleets.

INC-042AMR-17WH-EAST / aisle B414:32:18.02 UTCReconstructing

Investigation of incident 042: LiDAR degradation leads to localization instability, planner hold, and a protective stop on AMR-17.

Fragmented signals → correlated clock

  1. .04
  2. .21
  3. .47
  4. .63
  5. .71
  6. .89

Causal chain

Root cause

Localization instability caused by degraded LiDAR observations. Evidence traces from return density through pose covariance to the protective stop.

The problem

When autonomous systems fail, engineers become detectives.

A single incident can span bags, telemetry, cameras, LiDAR, IMU, localization, planner state, controller output, and logs. The data exists. Reconstructing causality still happens by hand.

Incident sources — uncorrelated

ROS / MCAP

bags · topics · tf

Telemetry

rates · estop · power

Cameras

streams · frames

LiDAR

returns · intensity

IMU

accel · gyro

Localization

pose · covariance

Planner

costmap · trajectory

Controller

cmd_vel · safety

System logs

nodes · diagnostics

Human investigation

An engineer opens each source in a different tool, aligns clocks by hand, and guesses which events belong to the same failure.

Root cause

Reached hours later — if the trail can be reconstructed at all.

The data exists. The intelligence to connect it doesn't.

Product

Observability tells you what happened.
Pharion investigates why.

Visualization and monitoring remain essential. Engineers still have to stitch timelines, compare robot state, and argue about causality. Pharion sits above that stack — it does not replace it.

  1. 01InfrastructureRobots, fleets, and the ROS / ROS 2 runtime already producing state.
  2. 02DataMCAP, bags, telemetry, LiDAR, cameras, IMU, GPS, logs, diagnostics.
  3. 03ObservabilityFoxglove, RViz, Grafana, and logs show what the system did. Necessary. Not an investigation.
04PHARIONInvestigation intelligence: correlate robot state, telemetry, logs, and sensor behavior into a causal chain.

Timeline

Reconstructed event clock

Evidence

Observable system behavior

Root cause

Why the failure occurred

Recommendation

What to change next

How it works

From fragmented streams to a defensible investigation.

  1. 01

    Ingest

    Connect the data autonomous systems already produce: ROS / ROS 2, MCAP, telemetry, cameras, LiDAR, IMU, GPS, diagnostics, and logs.

  2. 02

    Reconstruct

    Rebuild a single incident timeline across sensors, localization, planner state, controller state, and software events.

  3. 03

    Investigate

    Correlate anomalies and causal relationships — not just correlated charts, but ordered evidence.

  4. 04

    Explain

    Produce an evidence-backed root-cause analysis and recommended corrective actions the team can act on.

Product

An investigation, not a dashboard.

Incident #042 — AMR-17 stopped in aisle B4. The interface below is the kind of workspace Pharion is built to be: dense, ordered, and accountable to the data.

INC-042AMR-17Fleet: WH-EAST14:32:18.02 UTCprotective_stop

Robot state

LiDAR conf.

0.41

↓ 0.37

Pose cov.

0.18

↑ 0.12

Loc. quality

LOW

warn

Planner

HOLD

changed

Causal reconstruction

LiDAR degradation → localization confidence ↓

Pose covariance ↑ → planner receives unstable pose

Controller enters safety state → robot stops

Recommendation

Inspect LiDAR occlusion and return-quality thresholds on AMR-17. Raise localization covariance gate before planner execution. Replay incident against last known-good map at 14:31:40.

Explainability

Don't just get an answer. Get the evidence.

Every conclusion is traceable to observable system behavior. If it cannot be shown in the data, it does not ship as a finding.

Likely root cause

Localization instability caused by degraded LiDAR observations.

14:32:11.04

LiDAR confidence decreased

Return quality fell below 0.45 on consecutive sweeps.

14:32:11.21

Localization covariance increased

AMCL covariance crossed the operational bound.

14:32:11.47

Pose divergence detected

Estimated pose diverged from wheel odometry by 0.42 m.

14:32:11.63

Planner state changed

Planner froze the trajectory after receiving unstable pose.

14:32:11.71

Safety controller triggered

Protective stop issued; drive torque commanded to zero.

Incident timeline

The story of a failure, reconstructed.

Pharion aligns events from independent subsystems onto one clock so the sequence is visible — not inferred from memory.

  1. LiDARConfidence ↓ 0.78 → 0.41
  2. LocalizationCovariance ↑ 0.06 → 0.18
  3. State est.Pose divergence 0.42 m
  4. PlannerState HOLD (was TRACK)
  5. ControllerSafety trigger: protective_stop
  6. BaseRobot stopped · v = 0.00 m/s

Where it applies

Autonomous systems where failure investigation matters.

Pharion starts with industrial autonomy. The same investigation problem appears anywhere a machine acts without a human in the loop.

Initial wedge

Industrial robotics

Warehouse AMRs, autonomous forklifts, and manufacturing cells — where unexplained stops halt operations.

Horizon

Drones & defense

Autonomous aerial systems in dynamic environments, where after-action reconstruction cannot wait on tribal knowledge.

Horizon

Space robotics

Satellites, spacecraft, and remote platforms where replay is expensive and the system cannot be walked up to.

Horizon

Advanced robotics

Humanoids, field robots, inspection, and agricultural systems — high-DOF platforms with dense multimodal state.

Why now

The more autonomous systems scale, the more expensive unexplained failures become.

  1. 01

    Autonomous systems are scaling

    Robots are leaving controlled prototypes and entering live operations — warehouses, yards, fields, and airspace.

  2. 02

    Systems are becoming more complex

    Perception, planning, control, and learned models now share a vehicle. Failures cross subsystem boundaries.

  3. 03

    Reliability is a competitive advantage

    As fleets grow, unexplained downtime compounds. Teams that can investigate faster ship safer systems.

Vision

The reliability intelligence layer for autonomous systems.

Today

Engineers investigate failures manually — across bags, logs, and tribal knowledge.

Tomorrow

Autonomous systems continuously understand their own failures, with evidence attached.

Eventually

Every robot has a memory of what happened, why it happened, and what changed.

Team

Built by engineers who have lived the investigation.

Pharion is being built by people with experience across robotics, autonomy, AI, and the software systems that keep fleets running. Founder details will land here.

  • Robotics
  • Autonomy
  • AI systems
  • Engineering infrastructure

When your autonomous system fails, know why.

See how Pharion turns fragmented system data into an evidence-backed investigation.