Databricks cost attribution, down to the job, cluster, and team.
LakeSentry attributes Databricks spend automatically. It combines billing records, workload metadata, identities, tags, and your own attribution rules to map each cost line to a team, and labels every allocation with a confidence tier (exact, strong, estimated, or unattributed), so the number can be defended, not just reported.
Read-only by default. Write access only if you choose to execute actions.
Where Tag-Based Attribution Stops
Databricks gives you the raw material for attribution: system tables record every DBU, and tags travel with clusters and jobs where teams apply them. For dedicated, consistently tagged compute, that can be enough.
The gap opens on everything shared. All-purpose clusters and serverless SQL warehouses serve many users at once, so no single tag can own the bill. Platform overhead belongs to no job at all. And tagging discipline varies by team: the moment one pipeline ships untagged, someone is reconciling spreadsheets.
- Shared compute can’t be split by tags: a tag names one owner, the cluster served ten.
- Untagged workloads land in a bucket someone has to explain at month end.
- Every workspace reports in isolation; the account-level rollup is on you.
How LakeSentry Attributes Spend
A rules engine on top of a normalized cost ledger, built so every allocation can be explained.
Rules, mappings, and identity
Exact rules pin known resources to teams. Pattern rules match names, tags, and principal domains. Identity mappings resolve users and service principals to teams. Attribution works with whatever tag coverage you have today.
Shared compute, split by usage
Shared all-purpose clusters and serverless SQL are split by observed activity per user session, so the teams that used the compute carry its cost in proportion. Platform overhead distributes across teams by their compute-spend share.
Confidence tiers on every allocation
Every allocation is labeled exact, strong, estimated, or unattributed. Spend LakeSentry can’t defensibly assign stays visible at the workspace level instead of being smeared across teams. A detector warns you if that unattributed share starts growing.
System Tables and Tags vs LakeSentry
The native building blocks are solid. This is what changes when attribution runs on top of them.
| System tables + tags | LakeSentry | |
|---|---|---|
| Untagged spend | Lands in an unallocated bucket; manual research | Resolved through identity mappings and pattern rules; the remainder is flagged at the workspace level |
| Shared compute | One tag, many users — no fair split | Session-based split by observed usage per user |
| Platform overhead | Sits outside job-level views | Distributed proportionally by each team’s compute share |
| Cross-workspace view | Per-workspace queries you join yourself | One normalized ledger across every workspace |
| Auditability | Depends on who wrote the SQL | Each allocation carries its rule, path, and confidence tier |
Go Deeper
The mechanics live in the docs; the reasoning lives on the blog.
Cost attribution and confidence tiers
The full attribution model: rules, evaluation order, and confidence tiers.
FinOps 101 for Databricks
Why transparency comes before optimization — and where attribution sits on the maturity path.
Databricks native cost tools, compared
What system tables and dashboards cover natively, and where LakeSentry picks up.
Cost attribution ships in every tier, including Free: unlimited workspaces, one user, three months of history.
Free
$0€0
Standard
$499€499/mo
Pro
$849€849/mo
No per-DBU tax. No per-workspace fees. Paid tiers billed annually.
Frequently Asked Questions
Do we need to tag everything before LakeSentry can attribute costs?
How does LakeSentry attribute shared clusters and serverless SQL?
What happens to spend that can’t be attributed?
What data does LakeSentry read to do this?
How is this different from the Databricks usage dashboards?
See it in your own environment.
The free tier covers unlimited workspaces with three months of history. No card required.