AI/ML Engineer · research systems

Pavel
Kazantsev

I build experiments that can show their work.

I build AI products, evaluation systems and data-intensive experiments, backed by seven years of production software engineering. I publish the architecture, evidence and failures.

LLM integrations · NLP & RAG · ML pipelines · agentic workflows

Writing

Engineering what survives

View all
My Football Model Passed Validation. A Five-Check Audit Killed It.
Experiment 002 · Rejected5 min read

My Football Model Passed Validation. A Five-Check Audit Killed It.

Val away CLV was +152 bps, confidence interval excluded zero. Then I ran a five-check forensic audit: opening-line baseline, Brier skill, feature contamination, incremental CLV, and season stability. Four failed.

PythonMLLightGBMQuant researchReproducibility
Read the case study →

What I work on

AI beyond the demo

The model is one component. The product also needs data, orchestration, evaluation, interfaces and failure handling.

LLM integrations

OpenAI, Anthropic, Gemini, open-weight models via OpenRouter. Structured outputs, function calling, context management, cost controls and provider fallback routing.

NLP & RAG systems

Retrieval-augmented generation, embedding pipelines, reranking, chunking strategies and structured output extraction — built for production accuracy, not demo conditions.

ML pipelines & research

Walk-forward OOS, leakage controls, point-in-time data, cost models and immutable experiment artifacts. Results that can be interrogated, not just reported.

Agentic workflows

Typed tools, resumable stages, deterministic evaluation gates and explicit human approval — instead of autonomous systems with hidden state.