01 / GenAI & AI Evaluation
Evaluate whether AI systems actually reason.
PDS develops rigorous evaluation approaches for mathematical, scientific and multi-stage reasoning — with emphasis on reproducibility, failure modes and measurable performance.
REASONING
Mathematical & technical reasoning
Tasks designed to distinguish genuine problem solving from pattern matching.
LONG HORIZON
Multi-stage reasoning
Evaluation of systems that must maintain consistency across long chains of dependent steps.
AGENTS
Agent & workflow evaluation
Measure tool use, planning, recovery and outcome quality across complex workflows.
FAILURE ANALYSIS
Why did the model fail?
Structured diagnosis of numerical, logical, retrieval and instruction-following failures.