Skip to content
PHARION

Reliability intelligence

AI incident investigation for autonomous systems.

Pharion reconstructs failures across ROS, telemetry, sensors, and robot state - identifying what happened, why it happened, and how to prevent it.

Built for robotics engineering teams operating autonomous fleets.

INC-042AMR-17WH-EAST / aisle B414:32:18.02 UTCReconstructing

Investigation of incident 042: LiDAR degradation leads to localization instability, planner hold, and a protective stop on AMR-17.

Fragmented signals → correlated clock

  1. .04
  2. .21
  3. .47
  4. .63
  5. .71
  6. .89

Causal chain

Root cause

Localization instability caused by degraded LiDAR observations. Evidence traces from return density through pose covariance to the protective stop.

The problem

When autonomous systems fail, engineers become detectives.

A single incident can span bags, telemetry, cameras, LiDAR, IMU, localization, planner state, controller output, and logs. Reconstructing causality still happens by hand.

Incident sources - uncorrelated

ROS / MCAP

bags · topics · tf

Telemetry

rates · estop · power

Cameras

streams · frames

LiDAR

returns · intensity

IMU

accel · gyro

Localization

pose · covariance

Planner

costmap · trajectory

Controller

cmd_vel · safety

System logs

nodes · diagnostics

Human investigation

An engineer opens each source in a different tool, aligns clocks by hand, and guesses which events belong to the same failure.

Root cause

Reached hours later - if the trail can be reconstructed at all.

The data exists. The intelligence to connect it doesn't.

Product

Observability tells you what happened.
Pharion investigates why.

Visualization and monitoring remain essential. Engineers still have to stitch timelines, compare robot state, and argue about causality. Pharion sits above that stack - it does not replace it.

  1. 01InfrastructureRobots, fleets, and the ROS / ROS 2 runtime already producing state.
  2. 02DataMCAP, bags, telemetry, LiDAR, cameras, IMU, GPS, logs, diagnostics.
  3. 03ObservabilityFoxglove, RViz, Grafana, and logs show what the system did. Necessary. Not an investigation.
04PHARIONInvestigation intelligence: correlate robot state, telemetry, logs, and sensor behavior into a causal chain.

Timeline

Reconstructed event clock

Evidence

Observable system behavior

Root cause

Why the failure occurred

Recommendation

What to change next

How it works

Ingest. Reconstruct. Investigate. Explain.

Connect the streams the robot already emits, align them to one clock, find the causal chain, and ship a recommendation the team can act on.

  1. 01Ingest

    Connect the streams the robot already emits: ROS / ROS 2, MCAP, telemetry, cameras, LiDAR, IMU, GPS, diagnostics, logs.

    • MCAP / bagsamr17_1432.mcap
    • LiDAR/scan · intensity
    • Localizationamcl · pose cov.
    • Plannercostmap · traj
    • Controllercmd_vel · estop
    • Logsnodes · diagnostics
  2. 02Reconstruct

    Align sensors, localization, planner state, controller output, and software events onto one incident clock.

    14:32 shared clock

    1. 11.04LiDAR
    2. 11.21Loc
    3. 11.47Pose
    4. 11.71Safety
    5. 11.89Stop
  3. 03Investigate

    Detect anomalies and ordered causal links - not co-plotted charts, but which event caused the next.

    1. LiDAR density ↓

      loc confidence

    2. Pose covariance ↑

      planner hold

    3. Unstable pose

      safety stop

  4. 04Explain

    Emit an evidence-backed root cause and the corrective action the team can ship.

    Root cause

    LiDAR degradation → unstable pose → protective stop.

    Recommendation

    Raise return-density watchdog; hold planner on loc quality < 0.6.

Investigation

An investigation, not a dashboard.

Incident #042 - AMR-17 stopped in aisle B4. Timeline, robot state, evidence, and recommendation live in one workspace, each finding pointing at observable behavior.

INC-042AMR-17WH-EAST / aisle B414:32:18.02 UTCprotective_stopinvestigation

Workspace for incident 042: LiDAR degradation leads to localization instability, planner hold, and a protective stop. Focus or select a subsystem, chain step, or evidence item to see related events.

Event clock · 14:32

Evidence in this incident

Root cause

Localization instability caused by degraded LiDAR observations.

Recommendation

Raise the LiDAR return-density watchdog on AMR-17. Gate planner execution when localization quality drops below 0.6. Replay against the last known-good map at 14:31:40.

Request a Demo

A demo walks the same reconstruction on your fleet data.

Evidence

Don't just get an answer. Get the evidence.

Every conclusion traces to observable system behavior. If it cannot be shown in the data, it does not ship as a finding.

INC-042AMR-17E1–E5

Root cause

Localization instability caused by degraded LiDAR observations.

Incident timeline

The story of a failure, reconstructed.

Independent subsystems, one clock. From 14:32:11.04 to 11.89 the sequence is visible - not inferred from memory.

INC-042AMR-1714:32:11.04 - 14:32:11.890.85 s

Where it applies

Autonomous systems where failure investigation matters.

The same investigation problem exists anywhere a machine acts without a human in the loop. We are not locked to a first vertical.

  • Warehouse AMRs

    Unexplained stops in aisles, docks, and congestion.

  • Autonomous forklifts

    Load, lift, and path failures that halt a shift.

  • Manufacturing cells

    Motion, perception, and safety events on the line.

  • Drones & defense

    Aerial systems where after-action reconstruction cannot wait on tribal knowledge.

  • Space robotics

    Remote platforms where replay is expensive and nobody can walk up to the machine.

  • Advanced robotics

    Humanoids, field, inspection, and agricultural systems - dense state, high cost of an unexplained fault.

Why now

The more autonomous systems scale, the more expensive unexplained failures become.

  1. 01Systems are entering the real worldRobots leave prototypes and enter live operations - warehouses, yards, fields, and airspace - where failures have a cost.
  2. 02Sensors, software, and models share the vehiclePerception, planning, control, and learned components fail across subsystem boundaries. The trail is no longer in one log.
  3. 03Reliability is the advantageAs fleets grow, unexplained downtime compounds. Teams that reconstruct causality spend less time guessing.

Vision

The reliability intelligence layer for autonomous systems.

  1. TodayEngineers investigate failures by hand - bags, logs, telemetry, and tribal knowledge.
  2. TomorrowAutonomous systems continuously understand their own failures, with evidence attached.
  3. EventuallyEvery robot has a memory of what happened, why it happened, and what changed.

Team

Built across robotics, autonomy, and infrastructure.

Experience across robots in the field, autonomy stacks, and the software that keeps fleets running. About us · Contact us.

  • Robotics
  • Autonomy
  • AI systems
  • Software systems
  • Engineering infrastructure

When your autonomous system fails, know why.

Fragmented data becomes an evidence-backed investigation - what happened, why, and what to change.

  1. Logs
  2. Telemetry
  3. Sensors
  4. Robot state
  5. Investigation