Red Team Testing
Probe policy controls with adversarial prompts. This is not part of recorded-run scoring.
Deterministic quality gates
Score completed, persisted runs without invoking a model, chat, retriever, or tool.
grounded-v1
Durable results
Regression view
Observed cohorts
No evaluations for this project.
Recommendations are review records only. Nothing here changes the running model, prompt, policy, retrieval, or configuration.
No optimization candidates for this project.
Probe policy controls with adversarial prompts. This is not part of recorded-run scoring.
Exercise input handling independently from evaluation reports.