01 / GenAI & AI Evaluation

Evaluate whether AI systems actually reason.

PDS develops rigorous evaluation approaches for mathematical, scientific and multi-stage reasoning — with emphasis on reproducibility, failure modes and measurable performance.

REASONING

Mathematical & technical reasoning

Tasks designed to distinguish genuine problem solving from pattern matching.

LONG HORIZON

Multi-stage reasoning

Evaluation of systems that must maintain consistency across long chains of dependent steps.

AGENTS

Agent & workflow evaluation

Measure tool use, planning, recovery and outcome quality across complex workflows.

FAILURE ANALYSIS

Why did the model fail?

Structured diagnosis of numerical, logical, retrieval and instruction-following failures.