Databases 13 min read

Lance Stable Row IDs: Principles, Trade-offs, and How to Enable Them

This article explains Lance's Stable Row ID feature, which assigns persistent logical identifiers to rows so they survive compaction and updates, reducing index maintenance costs at the expense of extra mapping storage and lookup overhead.

Big Data Technology Tribe
Big Data Technology Tribe
Big Data Technology Tribe
Lance Stable Row IDs: Principles, Trade-offs, and How to Enable Them

1. Why Stable Row IDs Are Needed

Lance organizes data into fragments. A row's physical address is computed as row_address = (fragment_id << 32) | local_row_offset. For example, a row at fragment 3, offset 10 may move to fragment 8, offset 25 after compaction. If external references or indexes store physical addresses, they must be updated on every layout change.

Stable Row IDs introduce a logical identifier that remains constant while the physical location changes:

Before compaction: row_id 42 → fragment 3, offset 10
After compaction:  row_id 42 → fragment 8, offset 25

The article distinguishes three identifiers:

_rowaddr : Current physical address; may change after compaction.

_rowid (with Stable Row IDs enabled) : Logical row ID; remains unchanged after compaction.

_rowid (without Stable Row IDs) : Same as physical address; may change after compaction.

Without the feature, _rowid cannot be treated as a persistent row identifier.

2. Does “Stable” Include Updates?

The current Row ID specification states that ordinary updates preserve the logical row ID. For example:

Before update: row_id 42 → original data
After update:  row_id 42 → updated data

When an update creates a new physical row, the logical ID maps to the new address and the old physical row is marked deleted. The author notes a version difference: the earlier move-stable functionality only guaranteed ID stability across moves and compaction, not updates. Some API comments still reflect the old behavior and should not be applied to the current implementation. Also, “updating the same logical row” differs from “deleting and re-inserting a row”; the latter does not inherit the original ID.

3. How Stable Row IDs Reduce Index Maintenance Cost

If a vector index stores physical addresses, e.g., Xiao Ming's vector → (fragment 3, offset 10), compaction forces the index to handle the address change. Options include rewriting the index, using a Fragment Reuse Index to map old to new addresses, or letting affected index segments become stale.

If the index stores Stable Row IDs instead ( Xiao Ming's vector → row_id 42), the index entry stays valid. At query time, the logical ID is resolved to the current physical location. This avoids index writes caused solely by row relocation.

Crucial premise: the benefit only materializes when the index actually stores Stable Row IDs. Enabling the feature does not automatically confer the same advantage to all index types or versions. If the indexed data values change, normal index maintenance is still required.

4. Where the Row-ID-to-Physical-Address Mapping Lives

The mapping 42 → (fragment 8, offset 25) is a logical view. The persisted structure is a per-fragment sequence of row IDs ordered by physical row position. For example:

fragment 3: row_ids[10] = 42
fragment 8: row_ids[25] = 42

Given the fragment ID and the sequence, the physical address is derived. The sequences are compressed (ranges for consecutive IDs, other encodings for gaps). Three storage modes exist:

Regular Inline Storage

Compressed sequences are kept in fragment metadata and persisted with the manifest. This avoids extra file reads but the inline data is rewritten on every new manifest version.

Hidden Column Storage (Current Branch)

Larger sequences can be placed in a hidden _rowid column inside the data file to reduce manifest size. This requires explicit configuration and is not yet a released feature as of the article's writing.

In-Memory Reverse Lookup

At query time, Lance builds a RowIdIndex (row ID → current physical address) from the current fragment sequences and deletion information, and caches it. No full global reverse-mapping file is written on each compaction.

5. How Compaction Maintains the Mapping

For compaction paths that reorganize rows, the process is:

Read the row ID sequences of the old fragments being merged.

Remove IDs of deleted rows.

Preserve the surviving rows' original IDs.

Re-split the sequence according to the new fragments' row counts.

Commit the new sequences with the new fragments and new version.

Example: merging fragment 3 ( offsets 0,1,2 → row_ids 40,41,42) and fragment 4 ( offsets 0,1 → row_ids 43,44), with row_id 41 deleted, yields new fragment 8 with offsets 0,1,2,3 → row_ids 40,42,43,44. The new sequence implies row_id 42 → (fragment 8, offset 1). Old versions continue to interpret rows using their own metadata.

6. Mapping Maintenance Also Adds Write Overhead

Stable Row IDs do not eliminate maintenance cost. Extra costs include:

Writing new fragments' row ID sequences.

Rewriting inline metadata with each manifest version.

Writing hidden column data when that storage mode is used.

Building or loading the row ID lookup structure at query time.

Translating logical IDs to physical addresses when fetching data.

Potential savings come from two factors:

Location Information Can Be Compressed

One million consecutive IDs need not be stored as one million (row_id, fragment_id, offset) triples; range encoding can represent them compactly. However, frequent updates, deletes, or reordering can make sequences more complex, increasing storage and lookup cost.

Multiple Indexes Can Share the Location Information

If a table has three indexes (vector, text, scalar) all storing Stable Row IDs, they can share a single mapping layer:

Vector index ─┐
Text index  ─┼→ row_id → current physical address
Scalar index ─┘

After a row moves, only the shared mapping is updated, avoiding rewrites in each index. Actual gains depend on index structure, affected data volume, and index file rewrite granularity.

7. Why It Is Not Guaranteed to Be Faster

Stable Row IDs represent a trade-off:

May reduce: secondary index maintenance due to physical address changes
May increase: row ID sequence storage, mapping construction, query-time resolution

Physical-address indexes can also defer rewrites via Fragment Reuse Index. Therefore, one must compare concrete implementations and workloads; it is incorrect to assume

enable Stable Row IDs = compaction and queries are always faster

.

8. Default State and How to Enable

As of Lance v13.0.0, both Python write_dataset and Rust WriteParams default to disabled Stable Row IDs. To enable on a new dataset:

import lance
import pyarrow as pa

data = pa.table({"name": ["Alice", "Bob", "Carol"]})
dataset = lance.write_dataset(
    data,
    "./example.lance",
    enable_stable_row_ids=True,
)
assert dataset.has_stable_row_ids
table = dataset.to_table(with_row_id=True)

For existing datasets, passing enable_stable_row_ids=True on subsequent writes does not automatically migrate. Rust provides Dataset::migrate_to_stable_row_ids() for migration, which involves index and concurrent-write constraints; follow the documentation for the specific version.

References

Lance Row ID & Lineage Specification

Lance Index: Compaction & Address Remapping

Lance Index: Stable Row ID

Stable Row IDs Production-Readiness Project #8931

Early Move-Stable Row IDs Project #2307

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

migrationcompactiondatabase internalsLancevector indexStable Row IDdata layoutrow addressing
Big Data Technology Tribe
Written by

Big Data Technology Tribe

Focused on computer science and cutting‑edge tech, we distill complex knowledge into clear, actionable insights. We track tech evolution, share industry trends and deep analysis, helping you keep learning, boost your technical edge, and ride the digital wave forward.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.