physical intelligence · benchmarked

The benchmarking platform for physical intelligence.

pera helps teams building physical AI evaluate VLA and robot-policy models under reproducible benchmark protocols, with pinned tasks, seeds, scoring rules, episode videos, action traces, and verification-ready evidence.

pinned protocols per-episode evidence verification commands

how pera works

From physical-AI model submission to verifiable result.

Every pera evaluation follows a visible, reproducible workflow.

why pera is different

The score stays attached to the evidence.

pera is built for researchers, model builders, and evaluation agents expanding physical AI who need to inspect why a result happened, not just cite the aggregate number.

Contracts before execution

A policy declares cameras, state format, control rate, and action conventions. The resolver records a direct, adapted, or refused model-benchmark pairing before simulation.

Failures stay inspectable

Result pages expose task-level scores, failure labels, episode videos, review coverage, and whether a label was manually reviewed or inferred from a task pattern.

Verification has clear limits

Env-verified rows attest environment execution, declared protocol, scoring, transcripts, and evidence integrity. Signed model-identity attestations are a later tier.

Conflict-of-interest disclosure: PulseVLA is verapulse's own model. Its rows use the same protocols, evidence surface, and disclosure rules as the other models.

benchmarks and compatibility

Supported protocols, model rows, and evidence status.

Open a cell to inspect the full row. Protocol cards show the simulator, episode budget, step cap, task list, and number of evaluated models.

- models · - protocols · - published result rows

Select any score to inspect the full result page.

reproduction reports

What ordinary leaderboards miss.

The reports document exact replications, quarantined runs, convention bugs, and cross-model effects that are easy to hide inside a single score.

same rules, your model

Put your physical-intelligence model under the same evaluation protocol.

Connect a compatible policy, run a supported benchmark, inspect every episode, and publish a result another team can verify.

Need help wiring an adapter or interpreting a result? Join the pera Discord for community support.