CASE STUDY

How Arlequin.AI Standardized Model Packaging with KitOps

Arlequin.ai at a Glance

  • Company: Arlequin AI is a topological deep learning research lab and AI platform delivering analytics that convert massive volumes of raw, unstructured, multilingual data into strategic decisions in minutes rather than days, giving governments and enterprises a decisive operational edge.
  • Team: Full ML chain in-house: researchers (topological deep learning), data scientists, ML engineers, and MLOps
  • Stack: PyTorch training, vLLM inference (~90% of production), MLflow experiment tracking, LakeFS dataset management, OCI-backed registry
  • Challenge: No standard way to link a model version to its dataset, config, and documentation
  • Solution: KitOps ModelKits as the packaging and versioning standard for every model artifact
  • Results: Reproducible, auditable, immutable artifacts for every iteration; simple promotion from dev to staging to production; reduced technical debt from consolidating tooling

The Company

Arlequin AI analyzes enormous volumes of structured and unstructured data to produce market intelligence its customers can act on. They position their models as a deterministic and sovereign alternative to LLMs from the handful of frontier model providers.

To move quickly they have optimized their data science and operations teams: a research team working on topological deep learning, while data scientists, ML engineers, and MLOps work side-by-side throughout the development and operations of each model. Training runs on PyTorch across a mix of vision-language, multimodal, topological and traditional models. Inference is standardized on SGlang, which serves roughly 90% of the company's production workloads. Today Arlequin deploys to cloud and on-prem environments, with multi-cloud deployments on the roadmap.

That breadth is exactly what makes model handoffs hard. When researchers, data scientists, and engineers all touch the same model on its way to production, the packaging layer either holds everything together or becomes the place where things fall apart.

Before KitOps: Reinventing the Wheel

Before adopting KitOps, Arlequin's team did what most ML teams do: they improvised. Model weights were pulled directly from Hugging Face or baked straight into Docker containers.

However, as the model complexity grew and they scaled things started to break. This workflow struggled the moment the team tried to package anything beyond weights. There was no standard for tying a model version to the dataset it was trained on and the config it needed to run. Every workaround felt like rebuilding the same wheel.

The MLOps landscape is fragmented. Before KitOps and ModelPack, there was no clear standard for packaging a model with everything it actually needs.
Aymeric Alixe
MLOps/DevOps Engineer
Arlequin AI

Finding the Standard

The team discovered KitOps through the MLOps community, on Reddit, at conferences, and in Discord, before the project joined the CNCF.

They spoke with KitOps maintainers from Jozu early and recognized the approach immediately: package the model, its documentation, its dataset references, and its configs as a single versioned OCI artifact, and store it in the registry infrastructure teams already run.

Adoption at Arlequin took time, but split along familiar lines. The engineering team moved fast; already familiar with Docker and OCI concepts, which made ModelKits feel native. The research team took longer and needed some education on what OCI is and why it exists. Once that clicked, the whole chain was on the same standard.

It felt like discovering containerization. If you were suspicious and then used KitOps, you wonder how you could not have used it before. It feels like Docker for models.
Aymeric Alixe
MLOps/DevOps Engineer
Arlequin AI

How Arlequin Uses KitOps Today

ModelKits are now the packaging and versioning standard for Arlequin's models.

Each ModelKit carries the model card, a pointer back to the experiment in MLflow, references to the datasets managed in LakeFS, and the config files the model needs to run. Everything is stored in an OCI-backed registry alongside the rest of the company's artifacts.

That structure powers the workflows that matter in production. Model promotion across dev, staging, and production is a registry operation rather than a manual handoff. Shadow traffic and A/B testing run on promoted ModelKits, so the team always knows exactly which model version, trained on which data, with which config, is serving which slice of traffic.

Results

The wins the team is most excited about are the ones that compound:

  • Reproducibility. Any model in production can be traced back to its exact weights, dataset references, and configuration. This only gains value as versions increase.
  • Auditability and immutability. ModelKits are tamper-evident, versioned artifacts. What was tested is what ships. This makes troubleshooting simple and quick.
  • Simple promotion. Moving a model from dev to staging to production no longer requires reassembling context from three different tools, speeding development.
  • Less technical debt. Retiring the patchwork of blob storage, registry hacks, and container workarounds removed a whole class of maintenance work, allowing them to focus on operating their models at growing scale.

Precise time savings are hard to pin down, but the team estimates that packaging and versioning workarounds alone would have consumed one to two weeks per year. The bigger payoff is the elimination of a critical failure mode: nobody has to ask which version is actually running, or whether the dataset and config match the weights. Getting that wrong for the security sensitive customers Arlequin deals with could be business ending.

Advice for Other Teams

Use the standard. It's community-backed and it does the job well. Not using it is like refusing to use containerization for your code.
Aymeric Alixe
MLOps/DevOps Engineer
Arlequin AI

Get Started by
Speaking with our
Engineering Team

Connect with our team to start a conversation. We're ready to collaborate, troubleshoot, and help you move forward.

Let's talk
KitOps is an openly governed CNCF project and the largest implementation of the CNCF ModelPack specification, the open standard for packaging and versioning AI/ML projects as OCI artifacts. Contributors include engineers from Jozu, Red Hat, PayPal, ANT Group, and ByteDance. Get started at kitops.org.