TrustifAI/typed_evals: use Jev, a System One model, to judge generated responses and recorded agent
Typed Evals Evaluate responses. Guard actions. A Python toolkit for evaluating LLMs, RAG, and agents—with optional calibration against human labels. Quickstart · Agent evaluation · Tool guar...
Continue Reading