⚓ Eurybia monitors model drift over time and securizes model deployment with data validation
-
Updated
Mar 23, 2026 - Jupyter Notebook
⚓ Eurybia monitors model drift over time and securizes model deployment with data validation
A curated list of awesome open source tools and commercial products for monitoring data quality, monitoring model performance, and profiling data 🚀
In this repository, we will present techniques to detect covariate drift, and demonstrate how to incorporate your own custom drift detection algorithms and visualizations with SageMaker model monitor.
These are my notes of the Udacity Nanodegree Machine Learning DevOps Engineer.
資料科學的日常研究議題
Simulation, testing and comparison of state of the art Unsupervised Concept Drift Detectors used in a batch Machine Learning scenario.
PipeRoll Seismograph - daily witnessed behavioural-drift readings for LLM APIs: 39 models measured against their own past, deterministic grading, Rekor-witnessed. Readings, not ratings.
In this project, we illustrate how the Kolmogorov Smirnov (KS) statistical test works, and why it is commonly used in Machine Learning (ML), Deep Learning (DL) and Artificial Intelligence (AI).
Detect behavioural drift between LLM versions before you upgrade. Compare model responses, classify regressions, and generate migration reports with validated prompt patches.
Learn how to handle model drift and perform test-based model monitoring
Behavioural attestation for deployed language models: is the model you serve the model you validated? 1 KB/position fingerprints that catch silent sampling filters, quantisation and template bugs.
End-to-end MLOps platform for sovereign country-risk prediction with MLflow, DVC, Docker, AWS SageMaker, CI/CD and production drift monitoring.
Hey LLM, you okay? — pyramid-ordered LLM testing CLI for CI/CD. One YAML for every layer, LLM-as-a-judge gates, and A/B triage that tells prompt regressions from model drift.
The Taravangian Test for AI: catch silent model degradation before your users do. SPC-based reasoning-quality monitoring for Claude, GPT, Gemini, and Grok.
LLMs as production extraction infrastructure: rule-vs-LLM triage, validated structured outputs, cost-capped model routing, a measured precision/recall eval + vendor-drift detection, entity resolution, and an Airflow 3 ETL DAG. Runs fully offline.
ModelPulse helps maintain model reliability and performance by providing early warning signals for these issues, allowing teams to address them before they impact users significantly.
A reproducible benchmark for whether commercial LLMs silently drift over time. Deterministic graders, balanced tasks, every raw response in git.
An end-to-end MLOps pipeline for real-time payment fraud detection, featuring a LightGBM classifier, SHAP explainability, and continuous drift monitoring.
Continuous, model-version-bound verification of agent skills: run a skill's eval suite with and without the skill, emit hash-verified dated receipts with confidence bands, and diff them into drift reports.
"Past performance of machine learning model is no guarantee of future results." We call it "model drift" or "model decay". This repository will introduce various methods for detecting model drift.
To associate your repository with the model-drift topic, visit your repo's landing page and select "manage topics."