Applied AI, software and intelligent systems.Built around measurable operating outcomes.
Pilot, build and managed operations

LLM Evaluation Template

Representative cases, acceptance criteria and regression evidence.

Get the download link by email

Enter your details and PlanckCyber will email the PDF link. No spam, unsubscribe anytime.

How to use this resource

A benchmark should represent the actual workflow, users, data, exceptions and harms - not only general model performance.

Important: General technical template; evaluation results do not establish absolute accuracy, security or compliance.

Evaluation identity

6 guided items

Case taxonomy

6 guided items

Measures

5 guided items

Failure review

5 guided items