The AI execution studio. Models are commoditizing; the value is the layer between a raw model and real-world value. We make AI reliable (data & evaluation), useful (automation & agents), and understandable (a creative studio for AI companies) - and everything we make is demonstrable.
Our client work is confidential by default, so what we publish is the kind of work you can verify yourself:
| repository | what it is |
|---|---|
| ravnlab-eval-harness | The Plausible-Wrong Benchmark (RL-PWB-1) - 20 expert-anchored trap cases across law, medicine, finance, and engineering, with a zero-dependency grading harness and an Inspect AI adapter |
| pwmetrics | Reference metrics for confident-wrong failure: grounding, calibration (ECE), the overconfidence gap, refusal correctness |
| review-gate-mcp | A working MCP server that asks a human when it isn't sure - extraction with a human-review gate |
| confident-wrong-field-guide | The research note behind all of it: a cross-domain taxonomy of the answer that reads perfectly and is false, every figure source-labeled |
The standard everything is graded under is public: The RavnLab Evaluation Rubric, v1.0 · free tool: The Confident-Wrong Checklist · site: ravnlab.com