Contracts before execution
A policy declares cameras, state format, control rate, and action conventions. The resolver records a direct, adapted, or refused model-benchmark pairing before simulation.
physical intelligence · benchmarked
pera helps teams building physical AI evaluate VLA and robot-policy models under reproducible benchmark protocols, with pinned tasks, seeds, scoring rules, episode videos, action traces, and verification-ready evidence.
how pera works
Every pera evaluation follows a visible, reproducible workflow.
why pera is different
pera is built for researchers, model builders, and evaluation agents expanding physical AI who need to inspect why a result happened, not just cite the aggregate number.
A policy declares cameras, state format, control rate, and action conventions. The resolver records a direct, adapted, or refused model-benchmark pairing before simulation.
Result pages expose task-level scores, failure labels, episode videos, review coverage, and whether a label was manually reviewed or inferred from a task pattern.
Env-verified rows attest environment execution, declared protocol, scoring, transcripts, and evidence integrity. Signed model-identity attestations are a later tier.
Conflict-of-interest disclosure: PulseVLA is verapulse's own model. Its rows use the same protocols, evidence surface, and disclosure rules as the other models.
benchmarks and compatibility
Open a cell to inspect the full row. Protocol cards show the simulator, episode budget, step cap, task list, and number of evaluated models.
- models · - protocols · - published result rows
Select any score to inspect the full result page.
reproduction reports
The reports document exact replications, quarantined runs, convention bugs, and cross-model effects that are easy to hide inside a single score.
Episode-count parity also exposed a quaternion branch bug that dropped one task from 98% to 58%.
Read report → report no. 2 · rejected run 84.4%Episode evidence revealed corrupted observations in 52 of 98 failures. The clean rerun scored 95.2%.
Read report → report no. 4 · contract failures 3 bugsRotation, proprioception frame, and image orientation all produced clean execution and wrong behavior.
Read report → report no. 5 · cross-model finding 4 / 4The strongest task-level gain was 26 additional successes out of 50, with no direction reversals.
Read report → report no. 3 · honest baseline 27%A verification system earns trust by reproducing low scores and failure structure, not only flattering numbers.
Read report →same rules, your model
Connect a compatible policy, run a supported benchmark, inspect every episode, and publish a result another team can verify.
Need help wiring an adapter or interpreting a result? Join the pera Discord for community support.