Cost attribution No per-job tagging required

Databricks cost attribution, down to the job, cluster, and team.

LakeSentry attributes Databricks spend automatically. It combines billing records, workload metadata, identities, tags, and your own attribution rules to map each cost line to a team, and labels every allocation with a confidence tier (exact, strong, estimated, or unattributed), so the number can be defended, not just reported.

Start Free

Read-only by default. Write access only if you choose to execute actions.

app.lakesentry.io/attribution
Attribution
Team rollup · Last 30 days · All workspaces
94% attributed
Team Spend · Confidence
Data Engineering $18,420 exact
12 exact rules
ML Platform $11,067 strong
tag + identity map
BI & Analytics $8,204 estimated
session split, 4 users
Marketing Data $4,306 strong
domain rule
Platform overhead $2,868 estimated
proportional share
Unattributed $2,967 unattributed
held at workspace
Total · $47,832 across 4 workspaces Attribution-declining detector: quiet

Where Tag-Based Attribution Stops

Databricks gives you the raw material for attribution: system tables record every DBU, and tags travel with clusters and jobs where teams apply them. For dedicated, consistently tagged compute, that can be enough.

The gap opens on everything shared. All-purpose clusters and serverless SQL warehouses serve many users at once, so no single tag can own the bill. Platform overhead belongs to no job at all. And tagging discipline varies by team: the moment one pipeline ships untagged, someone is reconciling spreadsheets.

  • Shared compute can’t be split by tags: a tag names one owner, the cluster served ten.
  • Untagged workloads land in a bucket someone has to explain at month end.
  • Every workspace reports in isolation; the account-level rollup is on you.

How LakeSentry Attributes Spend

A rules engine on top of a normalized cost ledger, built so every allocation can be explained.

Rules, mappings, and identity

Exact rules pin known resources to teams. Pattern rules match names, tags, and principal domains. Identity mappings resolve users and service principals to teams. Attribution works with whatever tag coverage you have today.

Shared compute, split by usage

Shared all-purpose clusters and serverless SQL are split by observed activity per user session, so the teams that used the compute carry its cost in proportion. Platform overhead distributes across teams by their compute-spend share.

Confidence tiers on every allocation

Every allocation is labeled exact, strong, estimated, or unattributed. Spend LakeSentry can’t defensibly assign stays visible at the workspace level instead of being smeared across teams. A detector warns you if that unattributed share starts growing.

System Tables and Tags vs LakeSentry

The native building blocks are solid. This is what changes when attribution runs on top of them.

System tables + tags LakeSentry
Untagged spend Lands in an unallocated bucket; manual research Resolved through identity mappings and pattern rules; the remainder is flagged at the workspace level
Shared compute One tag, many users — no fair split Session-based split by observed usage per user
Platform overhead Sits outside job-level views Distributed proportionally by each team’s compute share
Cross-workspace view Per-workspace queries you join yourself One normalized ledger across every workspace
Auditability Depends on who wrote the SQL Each allocation carries its rule, path, and confidence tier

Cost attribution ships in every tier, including Free: unlimited workspaces, one user, three months of history.

Free

$0€0

Standard

$499€499/mo

Pro

$849€849/mo

Full pricing

No per-DBU tax. No per-workspace fees. Paid tiers billed annually.

Frequently Asked Questions

Do we need to tag everything before LakeSentry can attribute costs?
No. Tags help where they exist (pattern rules can use them), but attribution also draws on identity mappings, resource ownership, and rules you define. Most teams start by mapping their highest-spend users and adding exact rules for long-lived shared infrastructure.
How does LakeSentry attribute shared clusters and serverless SQL?
By observed activity. Usage is split across user sessions, and each user resolves to a team through mappings. The allocation is labeled “estimated” so everyone knows it’s a proportional split, not a direct charge.
What happens to spend that can’t be attributed?
It stays at the workspace level, labeled unattributed. In our view that’s the right trade-off: a defensible 85% beats a guessed 100%. A dedicated detector alerts you when the unattributed share starts climbing, so coverage improves instead of eroding.
What data does LakeSentry read to do this?
LakeSentry reads Databricks system tables — cost and usage metadata. It never accesses your business data, notebooks, or query results. The connection is a read-only service principal, and Unity Catalog is required.
How is this different from the Databricks usage dashboards?
Native dashboards show usage per workspace and per SKU; they answer “how much”. Attribution answers “whose, and why”: cross-workspace rollups by team, shared costs split by usage, and every allocation explained.

See it in your own environment.

The free tier covers unlimited workspaces with three months of history. No card required.

Start Free