I work on validation and audit of AI-based decision systems, with a focus on Large Language Models (LLMs).
My current work explores how AI systems behave when making decisions — not only whether an answer is correct, but whether the decision remains consistent, robust, traceable, and resistant to changes in conversational context.
My main areas of interest include:
- LLM decision validation
- AI / Model Risk
- Decision correctness and consistency
- Robustness and stability testing
- Framing sensitivity
- Human–LLM interaction risk
- Decision traceability
- Reproducible Python-based validation
My background combines physics, software engineering, telecommunications, business analytics, and data science.
Most projects in this profile are independent and pet projects, created to investigate practical technical problems rather than as commercial work for large organizations.
Projects focused on validating and auditing the decisions of LLM-based systems — not just building with LLMs, but testing whether their decisions are correct, consistent, and fair under scrutiny.
A compact validation framework for an LLM-based financial decision system.
Human ↔ LLM ↔ Decision
The framework goes beyond accuracy and evaluates decision behaviour across multiple dimensions, including:
- Decision correctness
- Consistency and stability
- Framing sensitivity
- User-position mirroring
- Belief reinforcement
- Decision traceability
Portfolio project · MIT License
Does an LLM reward how you write, not what you say? An audit of LLM decisions on synthetic social-benefit applications, where identical facts are rendered in different writing styles — formal, rude/sloppy, emotional, and Russian.
Three checks: accuracy by style against a code-computed reference, the price of style (shift in approval rate vs. a formal baseline), and whether the model's rationale discloses style as a reason when a decision changes.
Portfolio project · MIT License
Can an LLM revise a decision when new evidence warrants a reversal — while remaining stable when it doesn't? A sequential evidence-update experiment using synthetic incident-diagnosis cases, grounded in the belief-revision and anchoring literature.
Separates final decision accuracy from decision trajectory correctness: V1 found a 100% final accuracy but only 60% fully correct trajectories, showing the model sometimes reverses its decision a stage too early or too late even when it lands on the right answer.
Portfolio project · MIT License
Does an LLM know when to answer, when to ask a clarifying question, and when to challenge a false or unsupported premise? A black-box validation of this three-way decision boundary (ANSWER / ASK / CHALLENGE), grounded in recent clarification and ambiguity-recognition research.
V1 scored 78.6% on action selection (83.3% on the core ask-vs-answer boundary), with clear failure modes: material ambiguity and unsupported premises prove substantially harder than straightforward missing-information cases. V2 (clarification quality — does it ask for the right thing?) is planned next.
Portfolio project · MIT License
An independent model risk and model validation framework for a machine-learning credit risk model.
Covers discrimination, calibration, stability and segment performance, with explicit separation between model performance, validation evidence, and validation conclusions.
Portfolio project · MIT License
A lightweight OpenWrt traffic shaping solution providing per-MAC upload and download bandwidth control using Linux tc and ifb.
Includes a backend, LuCI interface and packages for embedded devices.
Independent systems / networking project
Home Assistant custom integration for monitoring OpenWrt routers via SSH.
Provides CPU temperature, load, RAM, VPN and disk metrics.
Home Assistant integration combining the native conversation agent with a local Ollama LLM for free-form responses.
No cloud API required · MIT License
Before moving into Data Science and AI validation, my work included:
- Software engineering and system integration
- Telecommunications and network management
- Linux / Unix systems
- Technical support and infrastructure
- Business analytics and revenue forecasting
- Mathematical modelling
This background continues to influence how I approach AI systems: as systems to be tested, measured and understood — not only as models to be trained.
| Project | Job Type | Status |
|---|---|---|
| NPD prediction | Predicting the next purchase to better understand customer behavior and optimize business decisions. | Complete |
| NP Multilabel prediction | Predicting a user's next order as a set of product categories using a fully connected neural network. | Complete |
| Project | Study Type | Status |
|---|---|---|
| Basic Python | Data validation and user behavior comparison using real Yandex.Music data. | Complete |
| Data preprocessing | Studying the impact of family status and children on loan repayment. | Complete |
| Exploratory Data Analysis | Real estate market analysis based on Yandex.Real Estate data. | Complete |
| Statistical Data Analysis | Customer behavior analysis for tariff optimization. | Complete |
| Composite Project - 1 | Identifying patterns that determine game success. | Complete |
| Introduction to Machine Learning | Building a classification model for tariff selection. | Complete |
| Supervised Learning | Predicting bank customer churn. | Complete |
| Machine Learning in Business | Profit and risk analysis for oil extraction regions. | Complete |
| Composite Project - 2 | Predicting gold recovery rates in industrial processes. | Complete |
| Linear Algebra | Data transformation for personal information protection. | Complete |
| Numerical Analysis | Car price prediction based on technical characteristics. | Complete |
| Time Series | Forecasting taxi demand at airports. | Complete |
| Machine Learning for Text | Toxic comment detection for moderation. | Complete |
| Computer Vision | Age prediction from facial images. | Complete |
| Diploma Project | Predicting customer churn for a telecom operator. | Complete |
| Project | Job Type | Status |
|---|---|---|
| DL & NLP – GeoNames | Geographical name normalization using GeoNames. | Complete |
| DL & CV – Music Genre Prediction | Music genre classification based on album cover images. | Complete |