Cross-layer intelligence platform

No migration required. Connects to your existing stack.

Every tool sees symptoms.NaviLake finds the cause — and the fix.

Your lakehouse is slower, more expensive, and less reliable than it should be. NaviLake reads metadata from every connected layer and traces every problem back to what actually caused it — and what to change.

Metadata-only · Read-only · No migration · You run the fix

Causal chain detected

4 systems correlated · 4 events in chain

Signals2/4 live
dbtSparkIcebergBigQuery
dbtRoot cause

fct_order_events — no distribution mode set on the model

SparksoonBad write pattern

Nightly job wrote 12,400 files averaging 8MB

IcebergsoonStorage degraded

File count up 340% — query planning 4x slower

BigQueryCost & latency impact

Reads 3.2x slower, +$610 this month across 9 dashboards

Impact, root cause & fixSee the full chain below ↓

One change in dbt. Four consequences across three systems. NaviLake connected them.

Every tool in your stack only sees its own layer.

Your warehouse tool knows what's expensive. Your catalog knows what exists. None of them talk to each other — so the engineer becomes the integration layer, manually stitching root causes together over and over.

×Single-layer tools

  • See cost, but not why
  • See file health, but not who's writing bad files
  • See lineage, but not whether it's actually used
  • 2–4 hours of manual stitching per incident

NaviLake

  • Sees cost and the upstream cause
  • Traces file health back to the model that caused it
  • Cross-references lineage with real usage
  • Correlated and ranked automatically

No migration

NaviLake reads your existing stack — BigQuery tables, dbt models, Iceberg snapshots, and Spark jobs as connectors come online. Nothing moves. Nothing changes in your environment.

Connect your way

Read-only OAuth by default — no software to install, no schema changes. Teams that require data to stay inside their perimeter can run NaviLake as a private collector inside their own environment.

Works on your existing stack

You don't need Iceberg, or Spark, or anything else to get value. Connect BigQuery. You're already finding problems.

From symptom to root cause

A BigQuery cost spike, traced across four layers back to a dbt configuration — impact quantified, action attached.

Causal chain detected

4 systems correlated · 4 events in chain

Signals2/4 live
dbtSparkIcebergBigQuery
dbtRoot cause

fct_order_events — no distribution mode set on the model

SparksoonBad write pattern

Nightly job wrote 12,400 files averaging 8MB

IcebergsoonStorage degraded

File count up 340% — query planning 4x slower

BigQueryCost & latency impact

Reads 3.2x slower, +$610 this month across 9 dashboards

Cost impact

$610 / month

Affected tables

1 degraded table

Pipelines at risk

9 dashboards affected

Governance gaps

1 config fix

Representative example. Layers marked soon join the chain as those connectors ship.

From connect to root cause — without moving anything.

Four read-only stages. Your data stays exactly where it is.

1

Connect

Read-only access to the tools you already run. No migration. No agents. Nothing moves.

2

Resolve

We link every connected asset — BigQuery tables, dbt models, Iceberg snapshots, Spark jobs — into a unified graph so cross-layer problems become visible.

3

Correlate

We trace problems across layer boundaries — from where they surface back to where they started.

4

Recommend

Each finding ships with the specific action to take. Your team decides when — and whether — to run it.

A job. A table. A pipeline. The whole estate.

Not just tables — the jobs that write them, the models that schedule them, and the queries that read them. Every finding says the same three things: what happened, what it costs, what to change.

A pipeline·Reliability

The data stopped arriving. Nobody noticed.

A table stopped receiving new writes. The pipelines reading it kept running on schedule — pulling numbers that had quietly frozen eleven days earlier. No error fired because the table still existed. The data just stopped changing.

Stale source, active consumers — detected before escalationdbtBigQuerylive
The whole estate·Governance

Nobody could say who owned a third of the data.

Not one table — hundreds. No model, no owner, no retention policy, several ingesting personal data by the hour. No alert ever fired, because nothing was looking across all of them at once.

340 tables surfaced in one viewBigQuerydbtlive
A table·Cost

Nobody had run maintenance in seven months.

Storage kept climbing while the data barely grew. Expired snapshots, orphaned files from failed writes, and metadata never rewritten had quietly taken a third of the footprint — taxing every query planned around them.

31% of storage reclaimedIcebergBigQuerysoon
A job·Performance

The job wasn't slow. It was fighting the table.

A nightly job spent most of its runtime shuffling data across the cluster. It grouped on one column while the table underneath was physically organised by another — so every run reshuffled the entire dataset just to reconcile the two.

4h 12m → 38m per runSparkIcebergsoon

The Spark tools see inside your jobs. The table-format tools see inside your tables. Your warehouse tool sees the bill. Not one of them can tell you the job is slow because of how the table is partitioned — because that fact doesn't live in any single layer. It only exists in the relationship between two of them.

From the live engine

What a finding looks like

Every finding has three parts: what's happening, why it matters, and exactly what to change. Your team decides when to act.

dbt + BigQueryLive

This model scans 12 million rows every run.

Context

analytics.daily_revenue_report rewrites from scratch daily — pulling all of warehouse.sales.orders each time.

Finding

The model has no incremental filter. Every execution bills for the full table regardless of how much data changed.

Action

Switch to incremental materialization with a date-range filter. Same results, 94% less compute.

$840 → ~$52/monthRule: DBT-004
BigQueryLive

This table costs more than it should — every day.

Context

warehouse.finance.transaction_log — 4.2 TB, queried daily by 6 pipelines, all without partition filters.

Finding

Every query scans the entire table. The access pattern is date-range based, but the table has no partitioning that supports it.

Action

Add date partitioning and update query filters. Alternatively, migrate to Iceberg on GCS — storage drops ~75%, query cost ~73%.

$1,224 → ~$329/monthRule: WH-007
Spark · Coming next

8 jobs shuffle the entire dataset — because of how the table is partitioned.

NaviLake will correlate Spark job metrics with table partition layout to find the exact mismatch causing the shuffle.

Iceberg · Coming next

Queries slowing down — not because of the query, but because of 40,000 small files underneath it.

NaviLake will surface snapshot bloat, orphan files, and metadata growth — with safe, sequenced compaction runbooks.

What's live today — and what's next.

Each new connector compounds the intelligence of every connector before it. Starting with the most common lakehouse stack — more warehouses, compute engines, and table formats follow the same pattern.

Your existing data stack

BigQuery

Warehouse

dbt

Transformation

Iceberg

Open table format

Spark

Compute

Intelligence layer

navilake.

Connect→Resolve→Correlate→Recommend

Root cause

Traced across layers

Quantified impact

Cost · performance · risk

Recommended fix

Evidence-backed · you execute

Live todayComing nextRoadmapAdditional warehouse · compute · table-format · storage systems
BigQueryLive

Usage, cost, query history, and query fingerprints — full-scan detection, cache analysis, write-stopped tables included.

dbtLive

Model-to-table mapping with confidence scoring, lineage, run cost attribution, full-scan detection, and schema drift.

IcebergComing next

Storage health — small files, snapshot bloat, delete ratios — with safe compaction and GC runbooks.

SparkComing next

Compute tuning — skew detection, disk spill, shuffle bottlenecks, and job-to-table attribution.

Built for your environment. Not inside it.

Metadata-only

We read schema, statistics, and job history — never your actual data files. No write access to your warehouse, ever.

You run the fix. We never do.

Every recommendation is a specific action for your team to review and execute. NaviLake never touches your environment.

A documented way back

Each recommended action comes with its rollback path written out, plus the evidence that produced it — frozen at the time of the finding.

Safe maintenance ordering

For Iceberg, we always sequence maintenance operations in the safe order. Running GC out of sequence has caused data corruption in production. We never let that happen.

Not a cost dashboard. Not a catalog.

Built by engineers who've spent years inside the data infrastructure teams that owned the problems NaviLake solves.

“Every data team we've ever been part of had the same failure mode: five excellent tools, each confident about its own layer, none of them agreeing on the full picture. We didn't want to build another excellent single-layer tool. We wanted to build the layer that sits above all of them — one that gets more valuable every time you connect something new to it, not less.”

— Founder, NaviLake

See your first causal chain in under 10 minutes.