Konrad Kowalski (rootsher)Principal Platform & Reliability Architect110011100100001111010010111000001100111110001011

Model Evaluation

posts (1)

  1. Part II: Standardizing AI in the SDLC Across the Organization9/10

    Evals and quality gates

    Observability tells you what the agent did, not whether it did it well. Evals next to the agent code work as regression tests for behavior and block rollout…