Benchmarks

The primary experiment is without skills vs with skills. We measure packaged knowledge, not which lab shipped the latest model.

Catalog stats

Knowledge uplift

+35.8%

Software Throughput

13.6×

Human intervention

5→0

Primary

Does giving Ruby-specific knowledge to an agent make it better at producing Ruby software?

Secondary

Does better knowledge produce more correct software per unit of time?

Relay

relay.example · same app, every run

A support inbox: tickets, assignment, and a Hotwire conversation thread.

  • Rails 8.1
  • Hotwire
  • PostgreSQL
  • RSpec

Protocol

Example run · August 2026 · GPT-5 on Relay

One prompt. One empty Rails app. Knowledge is the only variable — first none, then skills added in layers.

Build a small support inbox. Staff can list tickets, assign an owner, and reply in a live thread.

Full Skillfile

  • example/rails-conventions
  • example/secure-defaults
  • example/request-specs
  • example/hotwire-ui
  • example/job-queues

Without skills vs with skills

Ablation on GPT-5. Illustrative numbers, not a live lab.

Experiment Tests Time Iterations Tokens Human Throughput
Base model 31/44 24m 7 184k 5 4.9
+ Rails conventions 36/44 20m 5 151k 3 10.2
+ Testing knowledge 41/44 18m 3 126k 2 17.3
Full Skillfile 44/44 15m 2 103k 0 66.7