Versioned
Every edit is a proposal with a diff. Approve it, ship it, roll it back in one click.
Each point is a production prompt: versioned in git, tested on its own evals, scored on every version. Stop pasting prompts into code. Pull the one that passed.
Every edit is a proposal with a diff. Approve it, ship it, roll it back in one click.
Each prompt carries its own evals. They run on every change, before anything reaches production.
Every version is scored against its own evals. A number and a word. No vibes.
Rubricary is a public registry of versioned system prompts. Every prompt carries its own evals and receives a score on every version, so you can see which prompt actually passed before you use it.
Each version of a prompt runs against that prompt's own evals. The result is a number and a word: excellent (90 or above), good (75 to 89), fair (60 to 74) or poor (below 60). The score is tied to a specific model, so the same version can score differently on different models.
Evals are test cases written for one prompt: given these inputs, the output must meet these conditions. A version's score is how many of its cases pass. They live next to the prompt, in an evals file, and run on every change before anything reaches production.
Run one command: npx rubricary add <org>/<name>@<version>, for example npx rubricary add support/refund-request. It replaces the prompt text pasted into your code with a pinned, scored version you can update on purpose.
Prompts are versioned in git. Every edit is a proposal with a diff: approve it and it ships as a new version, or roll back to the previous one in one click. Each version keeps its own score, so you can compare them.
Yes. Yes. The validator at app.rubricary.com/validate lets you score a prompt of your own.
At app.rubricary.com/docs.
Your prompts are already in production. Put them under test.
Open the registry