Rubrkit
Grade, rewrite, and test your AI instructions
Rubrkit offers a robust evaluation system for various directives, including prompts, agent specifications, commands, skills, workflows, and rubr_flow procedures. Key capabilities include:
• 10-dimension rubric for comprehensive scoring
• Identification of weak instructions and areas for improvement
• Automatic generation of testable, rewritten instructions
• Version tracking for iterative refinement
• Multi-platform accessibility (Web, CLI, MCP)
This platform transforms vague requests into explicit, actionable directives. It provides an editorial loop, scoring clarity, context, constraints, output shape, and evaluation criteria. Users can paste an instruction, receive a detailed critique with a score, and then get a stronger, testable version with simple evaluations to prove its efficacy.
Rubrkit reviews diverse artifact types, applying a tailored 10-dimension rubric to match expected behavior for each. It focuses on highlighting weaknesses, explaining their impact, suggesting fixes, and validating whether these fixes hold up under real evaluation. For teams requiring strict control, rubr_flow offers an open, documented format for bounded, verifiable agent procedures with explicit work orders, bounded edits, and pass/fail verification.
Ideal for engineering and content teams, technical writers, or anyone developing precise directives that need to be consistently clear and perform reliably. It ensures instructions are not just generated, but are robust, testable, and reusable.