Quality · A/B · evidence gates · Databricks

Fabric Experiments

Quality thatships.

Measure, experiment, and gate AI quality on Databricks. Quality Center, A/B guardrails, and datasets/evals loops turn native evidence into a release decision.

Release evidence
Quality Center
Experimentation
A/B + guardrails
AI quality loops
Datasets + evals
MLflow · UC · SQL
Databricks native

Fabric ecosystem

DatabricksTemporal

One quality contract · many Databricks surfaces

Prove SQL, Delta, Lakeflow, Jobs, Unity Catalog, Apps, Model Serving, and edge experiments without rewriting the release decision.

SQLDeltaLakeflowJobsUnity CatalogMLflowApps

The 30 second answer

Databricks owns the platform. Experiments owns quality.

People often hear Fabric as a thin wrapper around MLflow or A/B tooling. That is the wrong layer. Unity Catalog, managed MLflow, Jobs, Lakeflow, Model Serving, and AI Gateway stay native. Fabric Experiments is the quality center, experimentation, and evidence-gate layer when measurement must become a release decision.

Databricks decides

  • Who may read or write data under Unity Catalog
  • Which models, serving endpoints, and MLflow runs are authoritative
  • Workspace identity, OBO, and service principals
  • Jobs, Lakeflow, SQL warehouses, Delta, and Apps execution

Fabric Experiments decides

  • How Quality Center aggregates and gates heterogeneous evidence
  • How A/B assignment, exposure, and statistical guardrails stay correlated
  • How datasets, evals, and live suites become promotion requirements
  • How tenant-scoped audit history explains what passed, failed, or aged out

Fabric does not recreate managed MLflow, Unity Catalog, Model Serving, AI Gateway, SQL, Delta, or Lakeflow. It drives and observes those native services, retains links to their authoritative evidence, and adds cross-workload release automation around them. Read the Databricks-native architecture.

When the quality layer earns its keep

Three production failures dashboards do not solve.

Lead with outcomes, not a catalog of Databricks services. Fabric Experiments differentiates after untrustworthy experiments, workspace-only failures, or AI quality that never became a gate.

The A/B result cannot be trusted

Traffic was assigned inconsistently, SRM went unnoticed, and the warehouse table does not match what the edge served. Growth and data teams argue about which number is real.

With Fabric Experiments. Signed manifests, deterministic assignment, exposure landing, and statistical guardrails keep variant delivery, analysis, and audit history on one evidence trail.

Green tests, broken workspace

Unit mocks pass while Unity Catalog permissions, Lakeflow refreshes, or Jobs fail under the identity that will run in production.

With Fabric Experiments. Local DuckDB contracts graduate to restricted-identity live suites against named Databricks targets, with cleanup, deep links, and publishable Quality evidence.

AI quality is a slide deck

Eval notebooks, ad-hoc judge scores, and chat screenshots never become a gate. A model or prompt ships because the demo looked good.

With Fabric Experiments. Datasets, eval suites, MLflow-linked runs, and Quality Center gates turn measurement into a promotion decision with tenant-scoped history.

The product boundary

Quality, evidence, and gates are the product boundary.

Notebooks, MLflow runs, and ad-hoc A/B tools already exist. Fabric Experiments starts where those stop: correlated measurement, promotion requirements, and a Quality Center decision operators can defend.

How it works

Experiment. Validate. Gate with proof.

A coherent contract from local quickstarts to live Databricks certification and Quality Center promotion.

  1. 01

    Experiment

    Design and deliver governed A/B tests with signed edge manifests, warehouse-backed exposure, and Studio analysis for frequentist, Bayesian, CUPED, and segment results.

  2. 02

    Validate and evaluate

    Describe expected behavior in Gherkin and TypeScript. Prove SQL, pipelines, jobs, permissions, and AI quality locally, then on real Databricks resources.

  3. 03

    Gate with evidence

    Publish named suites to Quality Center, require freshness and pass criteria, and promote configuration only when the release decision is defensible.

Choose the operating plane

Native where it matters. Connected everywhere else.

Fabric does not flatten Databricks into the lowest common denominator. It keeps MLflow, Unity Catalog, and compute authoritative while connecting experimentation, testing, and release gates around them.

Databricks · first class

Keep data and AI native. Add quality when shipping demands it.

Use managed MLflow, SQL, Jobs, and Model Serving as systems of record. Add Fabric when A/B, workload proof, and AI evals must share one promotion decision.

  • Native Unity Catalog and workspace identity
  • MLflow-linked evaluation evidence
  • Live packs for SQL, Delta, Lakeflow, Jobs, Apps
Explore Databricks deployment

Studio + edge · product surface

Design experiments, review results, enforce gates

Studio owns lifecycle, pipeline builder, Quality Center, and audit history. Edge delivery keeps assignment fast while warehouse analysis stays authoritative.

  • Hosted Studio for orgs and environments
  • Signed edge experiment manifests
  • CLI and CI evidence publishers
Explore Studio workflow

The Fabric difference

Quality, evidence, and gates are the product boundary.

A/B tools, notebooks, and MLflow already cover pieces of the loop. Fabric Experiments differentiates where measurement must survive assignment disputes, workspace identity, heterogeneous suites, and a production promotion clock.

Quality Center as the decision surface

Link native Databricks and MLflow evidence with BDD, live-suite, JUnit, and CLI results. Enforce freshness and pass requirements before production promotion.

Trustworthy experimentation

Assignment, exposure, conversion, statistical guardrails, and audit history stay correlated so A/B results remain reviewable under change.

Databricks-native proof

SQL, Delta, Lakeflow, Jobs, Unity Catalog, Apps, and Model Serving stay systems of record. Fabric drives them under restricted identity and captures deep links.

AI quality loops

Versioned datasets, deterministic judges, prompt hashes, traces, and warehouse evidence turn model and prompt changes into gated evaluation suites.

Local speed, live authority

DuckDB and offline contracts give seconds of feedback. The same intent certifies against a real workspace before Quality Center can open the gate.

Complete platform

Everything around the quality loop, in one coherent contract.

The differentiators are Quality Center, trustworthy experimentation, and gated evidence. The platform also includes the authoring, packaging, delivery, and operating surface needed to ship on Databricks teams.

Experimentation

Design, deliver, and analyze governed product and growth experiments.

  • YAML experiment manifests
  • Signed edge delivery
  • Web and Node SDKs
  • Exposure and conversion
  • Statistical guardrails
  • Studio analysis
Explore experimentation

Workload testing

Automate behavior for the Databricks surfaces data teams ship.

  • Gherkin + TypeScript
  • SQL and Delta packs
  • Lakeflow and Jobs
  • Unity Catalog
  • Apps and Lakebase
  • Restricted-identity live
Explore workload testing

AI quality

Evaluate applications against versioned datasets with governed evidence.

  • Quality Center
  • Datasets and evals
  • MLflow-linked runs
  • Model Serving judges
  • Prompt hashes
  • Token budgets
Explore ai quality

Governance

Turn heterogeneous evidence into a single release decision.

  • Production gates
  • Freshness requirements
  • Tenant-scoped audit
  • GitOps promotion
  • Enterprise controls
  • Evidence retention
Explore governance

Developer workflow

Ship from CLI, CI, and Studio without losing the evidence trail.

  • fx CLI
  • Local quickstarts
  • Target packs
  • JSON and JUnit evidence
  • Studio lifecycle
  • CI-friendly exits
Explore developer workflow

Delivery targets

Run the control plane and edge where your organization already operates.

  • Cloudflare Workers
  • Postgres control plane
  • Databricks worker
  • Temporal orchestration
  • Self-host static
  • Hosted Studio
Explore delivery targets

One promotion contract

Start with one suite. Keep the gate seams.

Describe experiments and quality requirements in ordinary YAML and TypeScript. The same suites publish into Quality Center whether they ran locally, in CI, or against a live Databricks workspace.

  • Named suites with freshness windows
  • SQL contracts, live workspace, and AI eval evidence
  • Promote only when every required gate passes
  • Deep links back to native Databricks and MLflow artifacts
See Quality Center gates
experiment.quality.yaml
# experiment.quality.yaml
experiment:
  key: checkout-rewrite
  traffic: 0.1
  variants: [control, treatment]

gates:
  - suite: sql-contracts
    maxAge: 24h
  - suite: agent-eval
    maxAge: 12h
  - suite: live-workspace
    maxAge: 48h

promote:
  require: all-pass
  evidence: quality-center

FAQ

Straight answers to the comparison questions

Use these when someone asks what Fabric Experiments provides that Databricks does not provide natively.

What does Fabric Experiments provide that Databricks does not provide natively?
Databricks owns MLflow, Unity Catalog, SQL, Delta, Jobs, Lakeflow, Model Serving, and workspace identity. Fabric Experiments owns the quality center, A/B experimentation, evidence gates, and datasets/evals loops that turn native signals into a promotion decision. See Databricks-native architecture.
Does Fabric replace MLflow, Unity Catalog, or Jobs?
No. Managed MLflow remains the evaluation system of record when you choose it. Unity Catalog, Jobs, Lakeflow, Model Serving, and AI Gateway stay native. Fabric drives and observes them, links their evidence, and adds cross-workload release automation.
Is this only an A/B testing tool?
No. Experimentation is one workflow. The same product validates Databricks workloads with Gherkin and TypeScript, evaluates AI applications, and gates production promotion in Quality Center.
Do I need Python?
No. Greenfield teams write step definitions, adapters, and support code in TypeScript. Python, SQL, R, and Scala notebooks, pipelines, and models can still be the workloads under test.
How do I start without a live workspace?
Follow the local quickstart or Databricks testing quickstart. Local DuckDB and offline contracts give fast feedback; certify a restricted identity on a real workspace before production gates.
Where does Studio fit?
Studio is the product surface for experiment lifecycle, results, Quality Center, and audit history. Open hosted Studio or read the Studio docs.
TypeScript quality platform · Fabric family

Start with one suite. Gate production when the evidence is real.

Run locally in minutes, then add live Databricks certification, A/B guardrails, and Quality Center promotion without rewriting the quality contract.

Fabric Experiments is built and supported by TechFabric.

Contact TechFabric