Who, what, when, where, why: On June 3, 2026 the OpenLineage project published OpenLineage 2.0 and announced first‑party integration commitments from Databricks, Snowflake and Google Cloud. The update formalizes a richer provenance schema and new APIs aimed at making training‑data lineage and dataset immutability easier to capture and share across platforms. For data engineers building machine‑learning pipelines, the release raises both an opportunity — standardized lineage across your stack — and a set of practical migration choices.

Why this matters now

Over the past three years, regulators and auditors have increasingly asked for demonstrable provenance of datasets used to train machine‑learning models. Meanwhile, enterprises are trying to scale ML while avoiding expensive post‑hoc forensic exercises. OpenLineage 2.0 aims to reduce friction by standardizing how systems describe dataset creation, transformation, access controls and immutability markers.

Vendor support from major cloud platforms shortens the path to adoption: integrations mean fewer bespoke connectors and less custom metadata plumbing. For teams responsible for model audits or compliance with corporate ML governance, that can translate to lower engineering overhead and faster time to evidence.

What’s new in OpenLineage 2.0 (key points)

  • Expanded provenance model: A lineage graph that encodes not just job‑to‑dataset edges but row‑level and column‑level transformations where feasible, plus dataset versioning identifiers.
  • Standardized dataset descriptors: Required and optional fields for dataset immutability (e.g., checksum, storage URI, retention policy), licensing and sensitivity labels.
  • Authentication and secure transport: OAuth/JWT profiles and signing patterns for metadata messages to support cross‑cloud evidence sharing.
  • Pluggable hooks for ML workflows: Lightweight SDKs and runtime hooks for Spark, Flink, dbt and popular orchestration tools to emit lineage events with minimal code.

Vendor commitments and scope

At launch, Databricks, Snowflake and Google Cloud published roadmaps that commit either to built‑in emission of OpenLineage 2.0 events or to first‑party ingestion connectors:

  • Databricks: Aiming for GA of OpenLineage 2.0 event emission from Delta Live Tables and Unity Catalog integrations in Q4 2026.
  • Snowflake: Announced a private preview of a Snowflake Metadata Connector that will consume OpenLineage 2.0 events and map them to Snowflake object versions in H2 2026.
  • Google Cloud: Committed to an integration between BigQuery Audit Logs and OpenLineage 2.0 metadata ingestion by end of 2026.

These commitments were documented in the project release notes and individual vendor blog posts published the week of June 3, 2026.

Impact for data and analytics engineers

Immediate implications are practical and technical:

  1. Metadata storage and cost considerations: OpenLineage 2.0 events can be verbose — include checksums, schema diffs and access policy pointers — so teams need to budget for increased metadata volume. Expect discrete storage needs (object store or dedicated lineage DB) and retention policies aligned with compliance requirements.
  2. Instrumentation work: Even with vendor integrations, pipeline instrumentation is not zero effort. Teams should plan for SDK upgrades, testing in CI, and validating that emitted lineage corresponds to actual data versions.
  3. Operationalizing lineage for audits: Lineage is only valuable when discoverable and trustworthy. Engineers must add QA checks (e.g., checksum validation, replay tests) and guardrails (immutable dataset tags) so evidence stands up to review.
  4. Security and privacy: Lineage metadata may contain sensitive URIs or schema details. Teams will need to apply metadata access controls and consider redaction or encryption in transit and at rest.

Concrete migration steps (a checklist)

Think of this as a pragmatic rollout plan you can follow in the coming months.

  • Inventory: Catalogue current lineage emitters, metadata stores and who consumes lineage (auditors, MLops, data catalog).
  • Pilot: Pick a representative pipeline (one batch ETL and one streaming pipeline). Enable OpenLineage 2.0 SDKs and route events to a staging lineage store.
  • Validate: Run checksum and schema‑match tests to confirm lineage correctness. Compare old lineage views with new OpenLineage graphs.
  • Secure: Apply RBAC to the lineage store, enable transport signing, and archive PD/PII fields per policy.
  • Automate: Add lineage‑emission checks to CI, and set alerts for missing lineage or unchanged dataset versions after transformations.
  • Document: Update runbooks, audit templates and onboarding docs so SREs and auditors know where to find and interpret lineage evidence.

What to watch out for

Don't assume vendor integrations immediately remove all gaps. Key risks:

  • Partial coverage: Vendor connectors often cover platform‑native jobs but not third‑party SaaS sources, custom microservices or edge data capture. You’ll need custom emitters there.
  • Event semantics mismatch: Different systems may interpret versioning or "update" semantics differently. Map semantics explicitly during the pilot phase.
  • Cost creep: High‑cardinality lineage (row‑level events) can explode storage and query costs. Use sampling or summarize at appropriate granularity.

Reactions from the field

Enterprise ML teams contacted for this story welcomed standardization but emphasized implementation work. A senior data-platform engineer at a Fortune 100 retailer (speaking on condition of anonymity) said: “Having a common schema will save months of custom mapping work, but we still need to decide which datasets require row‑level provenance and which can be coarse‑grained.”

Open-source contributors framed the release as evolutionary. Release notes describe 2.0 as focused on interoperability: “This version codifies lessons learned about dataset immutability and audit evidence,” the project notes read.

What's next

Expect a staged adoption throughout 2026 and into 2027. Practical timelines to watch:

  • Q3–Q4 2026: Vendor private previews expand; community SDKs (Python, Java, Scala) reach 2.0 compatibility.
  • H1 2027: Broader GA integrations in major cloud services and first corporate case studies showing reduced audit times.
  • Through 2027–2028: Third‑party SaaS vendors and orchestration frameworks (Airflow, Prefect, Dagster) finalize first‑class support.

FAQ: Common questions data teams ask

Do I need full row‑level lineage for every dataset?

No. Treat row‑level lineage as a targeted control for high‑risk datasets (PII, regulated training data, financial records). For many ETL feeds, dataset‑level or partition‑level lineage is sufficient and far less costly to maintain.

Will OpenLineage 2.0 replace my current data catalog?

Not directly. OpenLineage is a metadata exchange standard; it complements catalogs by providing richer runtime provenance. Most teams will route OpenLineage events into their existing catalog or use it to enrich catalog entries.

How should I handle sensitive information in lineage events?

Apply the same data governance controls as you do for business data: minimize sensitive fields, use tokenization or redaction in emitted metadata, and enforce RBAC on the lineage store and APIs.

How long before vendors fully support my stack?

Vendor roadmaps indicate staged support through late 2026 and 2027. For heterogeneous stacks, expect to do custom instrumentation for some sources in the short term.

What metrics should I track during rollout?

Track the percentage of production pipelines emitting valid 2.0 events, average latency between job completion and lineage availability, metadata storage growth, and the number of audit requests resolved using lineage evidence.

Standardization of provenance through OpenLineage 2.0 reduces the engineering tax of explaining where training data came from — but it does not eliminate the need for thoughtful instrumentation, storage planning and governance. For data engineers, the next six months are a good window to pilot, refine and bake lineage into CI and operational workflows so audits and model governance stop being reactive fires and become routine checks.