Migration and rollback¶
migrate is the only rebasis command that writes to your index. Everything on
this page exists because of that.
Status: experimental¶
migrate works, and every guarantee below is covered by a test that runs against
a real store on every commit. It is marked experimental for one reason: the
evidence behind it is a test suite, not a fleet of production indexes. Nobody has
yet run it against a million-record index they could not rebuild.
What that means in practice:
- Take a backup you can restore without rebasis. Not
--keep-original— that is rebasis' own shadow copy, and it protects against a bad adapter, not against a bug in rebasis. Copy the directory, snapshot the volume, export the collection. Whatever your store offers. - Migrate a slice first.
--limit 5000on a real collection, then look at what it did.rollbackis one command away. - Read the backend table below. Support is not uniform, and the difference is in what has been tested, not in what the code claims.
The other commands — probe, fit, eval — never write, and carry no such
caveat.
Backends¶
Each of these runs the full migrate-and-rollback suite on every commit: the
vectors actually change, the record count does not, text survives, and rollback
restores the originals byte for byte. "The originals" means something narrower
on a store that keeps compressed codes rather than vectors — If your index is
stored quantized, below, is what rebasis detects and what it then says.
| Backend | Migrate | Notes |
|---|---|---|
chroma |
tested | |
lancedb |
tested | |
qdrant |
tested | Local mode holds an exclusive lock on its folder; rebasis releases it on close(). |
sqlite-vec |
tested | |
faiss |
tested | Needs an IndexIDMap2 and a .meta.json sidecar. A write reorders the index, so the sidecar is rewritten with it. |
memory |
tested | Not persistent; used by the durability tests. |
Scale is the untested dimension, not the backend list. The suite runs on hundreds of records, not millions.
What it guarantees¶
It only upserts. It replaces vectors in existing records. It never deletes a record, never removes a payload field, never drops metadata.
It copies before it overwrites. Each batch's original vectors go to a shadow store before the new ones are written. A crash before the write costs nothing; a crash after the write, without the shadow, would cost the originals.
The queue is the checkpoint. Interrupt it — Ctrl-C, a closed laptop, a
kill -9 — and rebasis resume <job-id> continues from the last completed batch.
There is no separate checkpoint file to get out of sync with reality.
It verifies what it wrote. After each batch, a sample of the written records is read back from the store and compared. A store that silently fails to write is the most common source of silent data loss, and this is the check that catches it.
It checks again on a fresh connection at the end. The per-batch read-back
goes through the handle that did the writing, and a client that caches will
happily hand back what it has in memory rather than what reached disk. So when
the queue empties, migrate reopens the store from its URI and re-checks a
sample. That is the difference between "the store accepted it" and "the store
kept it".
Running it¶
migrate needs an adapter pointing out of the index, not into it. That is
the one thing it cannot be run without, and until this release nothing produced
one — so the documented sequence handed it the query map instead, and every guard
the tool had let that through. The reasoning is in what changed,
below; the short version is that an index rewritten with a query map answers
recall@1 0.000 to every query type there is.
# The map migrate needs. Note --direction; without it `fit` produces the
# query-side map, which `Bridge` serves with and `migrate` now refuses.
rebasis fit \
--store chroma:///path/to/db#documents \
--old <old-model> --new <new-model> \
--direction old_to_new \
--out forward.rbs
rebasis migrate --adapter forward.rbs --store chroma:///path/to/db#documents
It shows what it will do — how many records, which store, which adapter, which
job id, whether rollback is available — then a disk-space plan: the shadow copy,
the checkpoint and state, and the free space needed with a safety margin. If the
disk cannot take it, it stops there rather than filling the disk halfway through
and taking the shadow copy with it. Then it asks before starting; --yes skips
the question for scripts.
It holds the exclusive state lock for the whole run, so a second migrate,
rollback or gc --apply against the same state directory is refused rather
than interleaved. rebasis status and a bare gc take no lock.
Useful flags:
| Flag | What it does |
|---|---|
--batch 256 |
Records per batch. Adjusted automatically under memory pressure. |
--limit 5000 |
Stop after this many. Migrate an evening at a time. |
--max-memory 2GB |
A ceiling. The batch size is computed from it — you should not have to. |
--priority access --access-log log.jsonl |
Migrate what you actually read first, so quality improves where you will notice. |
--power-aware/--no-power-aware |
Pause on battery. On by default. |
--shadow-precision float16 |
Half the shadow's disk, a rollback that is close rather than exact. See below. |
--refit |
Refit the adapter part-way through, on records not yet migrated. Off by default; see below. |
--resume <job-id> |
Continue an interrupted job. rebasis resume <job-id> is the same thing. |
Whether to run it at all¶
Measured across 51 runs on seventeen corpora with human relevance judgements (the band):
| a completed migration delivers | 0.727 of a full reindex, on average |
| bridging, on the same runs | 0.719 |
| the two track each other at | Spearman 0.993 |
| migrating beat leaving the index alone in | 5 of 51 |
So migrating and bridging are worth the same amount, and both are usually worth less than doing nothing. That is ADR 10 reaching the document side: the same source space under the same family of map carries the same amount, whichever end you apply it to.
What migrating buys is the adapter leaving the query path — no map on the hot
path, no .rbs to ship with your service, and the new model querying its own
space. What it costs is rewriting every vector, the shadow copy behind it, and a
window in which the index holds two spaces. Choose on those grounds; quality is
not one of them.
What changed, and why¶
An adapter has a direction, and the two are mirror images:
| direction | maps | used by |
|---|---|---|
query_to_old |
a new-model query into the index | rebasis.Bridge |
old_to_new |
the stored vectors into the new model's space | rebasis migrate |
rebasis fit produced only the first, migrate never checked, and the README
showed one being piped into the other. Applying the query map to document vectors
passed every guard: the write landed, the count held, the text survived, the
read-back compared what was written against what came back, and the index-health
check measured the store's search against exact kNN over the vectors it now held.
None of those asks whether the vectors still mean anything.
Measured on data where both spaces are known exactly and the bridge itself scores recall@1 1.000 against the untouched index, the index a completed migration left behind answered 0.000 to a raw new-model query, 0.000 to a bridged query and 0.000 to an old-model query. Where the two models have the same width it failed silently; where they differ it failed with a dimension error at query time, which is why it survived on some ladders and not others.
Both directions are now guarded. migrate refuses a query_to_old adapter before
it opens the store, and Bridge.load refuses an old_to_new one.
Stopping short leaves two spaces in one index¶
--limit, --priority access and every pause all do the same thing to the
collection: for as long as the job is unfinished, some records hold the new
model's vectors and some still hold the old model's, and there is no query
that is correct against both.
bridge.to_index_space(q) correct for the records that have not moved
f_new(q) correct for the records that have
Whichever you send, part of the corpus is scored against a geometry it is not in. Nothing raises. The record count is right, the text is right, the ranking is wrong — and on a graph index the traversal itself is running over the mixture, so the damage is not limited to the scores of the records that moved.
This is the reason --limit is recommended above as a way to try a migration
rather than as a way to pace one. It is safe for the data — the shadow copy is
intact either way — and it is not safe for queries in the window before the job
finishes.
rebasis will not let that window be silent. It is named three times:
- in
migrate's preview, before you confirm, whenever--limitwill stop the run short; - at the end of any run that did stop short;
- by
rebasis status, unprompted, until the job finishes or is rolled back — including in--json, asmixed_space, so a script can refuse to serve.
$ rebasis status
...
This index holds two embedding spaces. Search results are not correct until
the migration finishes or is rolled back.
chroma:///path/to/db#documents holds two embedding spaces: 5,000 of 48,000
records (10%) have the new model's vectors and 43,000 still have the old
model's. …
rebasis migrate --resume job-8f2a1c4e0b73 (finish it)
rebasis rollback job-8f2a1c4e0b73 (put the index back)
Searching one anyway¶
Finishing the job or rolling it back are the two ways to resolve a mixed index. There is a third thing you can do while it is mixed, which is search it correctly:
from rebasis.serve import Bridge, MixedSpaceSearch
bridge = Bridge.load("adapters/minilm-to-bge.rbs")
with MixedSpaceSearch(store, bridge, job_id="job-8f2a1c4e0b73") as search:
hits = search.search(new_model.encode(["how do I deploy?"])[0], k=10)
It sends both queries and keeps only the half each one is right about — the
bridged query for the records that have not moved, the raw new-model query for
the ones that have — then merges the two through the isotonic calibrator in the
.rbs, which is what makes scores from two spaces comparable at all. Without a
calibrator it falls back to reciprocal rank fusion, which discards the scores
and uses ranks: strictly less information, and correct, where comparing raw
scores across two spaces is not.
Which records have moved is read from the manifest, not from the store.
rebasis does not write a rebasis_space field into your payloads; the whole
store contract is one write path that only ever replaces vectors, and the
migration queue already knows what it moved.
The cost is over-fetching: each side asks deeper than k and discards what
belongs to the other, scaled by how far the migration has got.
search.over_fetch reports it after every query, and the cheapest way to bring
it down is to finish the job.
Watching it¶
Takes no lock, so it works while a migration is running — which is exactly when you want it.
Refitting part-way through¶
Off by default. It samples records that have not been migrated yet, re-embeds them with the new model — the one recorded in the adapter's manifest, so there is no second flag to get wrong — refits on those pairs alone, and adopts the result only if it beats the adapter in use by 0.01 on a held-out slice.
It is for one situation. Measured over 216 cells, on a corpus that has not changed a refit is a pair-count effect worth a median +0.0075 at three times the fit budget, and nothing clears the adoption threshold. On an index that grew into a domain the adapter never saw — a vault that gained a department while the migration was running — 1,000 pairs drawn from what is left are worth a median +0.16 nDCG and beat 8,000 pairs drawn from the migrated half by +0.20. The numbers.
The guard is why it is safe to leave on in either case: it declines in the first and adopts in the second, and every attempt is audited with its reason.
An adopted adapter is written to .rebasis/adapters/<job>-refit-<n>.rbs and the
job is pointed at it, so rebasis resume continues with it rather than
reloading the file the job started with. It needs a store that returns document
text — there is no way to make a real pair without re-embedding one — and it
says so at the start of the run rather than at the first checkpoint.
Stopping it, and starting it again¶
rebasis pause <job-id> # stops after the current batch
rebasis resume <job-id> # picks up where it stopped
Killing the process is safe and always was: the queue is the checkpoint, and a shadow copy is written before the vector it copies is overwritten. What it is not is clean. A kill lands in the middle of a batch, and that batch's records are left in the store without having been read back and compared — verified writes are a per-batch guarantee, and half a batch does not get one.
pause stops at a batch boundary instead. It returns immediately and the job
stops a moment later, at the end of the batch it is in.
Like status, it takes no lock — the migration holds the state lock for its
whole run, so waiting for the lock would mean waiting for the thing you are
trying to stop. What makes that safe is that pause writes one column nothing
else writes. It records a request; only the engine ever says what state a job
is in. While the request is outstanding, status shows the job as
running (pausing), and --json carries it as a separate pause_requested
field.
A request never outlives the run it was made for: it is cleared when a run ends
and again when one starts. A process killed between rebasis pause and the
engine reading it cannot leave behind a flag that silently pauses tomorrow's run.
A paused job is a job stopped short, so everything under Stopping short leaves two spaces in one index applies to it. Pausing is not a way to stop safely and walk away; it is a way to stop safely and come back.
When your orchestrator stops it¶
A Kubernetes Job, an Airflow task and an Argo step all end a process the same
way: SIGTERM, a grace period, then SIGKILL. migrate catches the first
SIGTERM and turns it into exactly the request rebasis pause makes — the run
stops at the next batch boundary and says which signal asked. rebasis resume
picks it up.
The grace period has to outlast a batch. Kubernetes' default is thirty
seconds; a batch that takes longer is still killed part-way. Raise
terminationGracePeriodSeconds above your batch duration, or lower --batch
below the period you have. Running it in production
has the details.
The second signal is not caught, so a supervisor escalating — or a second Ctrl-C — stops the process at once.
Half the shadow, if you want it¶
The shadow is N x d x 4 bytes at the default float32, and half that at
float16. What the half costs is the bit-identical restore.
Measured over 68 corpus/model runs: no vector leaves the format, the top-10 set comes back on 99.78% of queries at worst, and nDCG@10 against human judgements moves by at most 0.0017 — inside ARR's own confidence interval. What does move is the order within the top ten, on about 2% of queries. The numbers.
float32 stays the default. A couple of gigabytes of temporary disk is cheaper
than any argument about whether 0.0017 mattered on your index, and the space
comes back when rebasis gc removes the shadow.
If you do use it, nothing pretends otherwise: the disk-space plan before the
confirmation says the rollback becomes approximate, and rebasis rollback prints
the precision it is restoring from — read off the shadow file itself rather than
off the job's configuration, because the shadow is the thing being restored from.
Rolling back¶
Restores the original vectors from the shadow copy. The shadow is written at float32, the same precision the vectors were read at, so what is kept is bit-identical to what was there.
What is restored is that, written back through the store's own upsert — which
is where the guarantee stops being rebasis'. A store that keeps what it is given
returns the same bytes. A store that normalises on write does not: Chroma in
hnsw:space=cosine shifts a vector by about 3e-08, and it does that to a plain
write-and-read with rebasis nowhere in it. The same collection at l2
round-trips exactly.
In practice that is a cosine similarity of 1.000000 against the original. It is not quite nothing: measured over three migrate-and-rollback rounds on a 5,183 document Chroma collection, recall@10 for one query in 300 moved, because a result sitting on the k boundary changed sides. Only ties that close can move at all — a shift of 3e-08 cannot reorder anything genuinely separated.
It is reported here because "bit for bit" is a claim somebody will check, and on a normalising store they would find it false.
The job records which store it wrote to, so you do not have to remember the URI months later — which is when a rollback is actually wanted.
If your index is stored quantized¶
Everything above assumes the store keeps what it is given. Increasingly it does not: int8 scalar quantization, product quantization and binary codes are how a large index is made to fit, and a store that holds a code cannot hand back the vector the code was made from.
migrate checks before it writes anything and says so in the plan, above the
confirmation. It does not refuse. A quantized index is a deliberate engineering
choice and migrating one is a legitimate thing to want; what you would not have
without the check is a correct reading of the paragraph above.
sqlite-vec is the exception, and the exception is the storage engine's. A
vec0 column declares an element type, and measured against the shipped
extension, inserting a float32 vector into an int8 or bit column is refused
outright — as is querying one with a float32 vector. rebasis produces float32 and
nothing else, so on such a table migrate and rebasis.Bridge were never going
to work. The backend declares can_upsert_vectors=False for them, which stops
migrate at the capability check before a job is opened rather than at a SQL
error after the first batch's originals have been deleted.
probe still works on an int8 table: one byte per component decodes to a
direction, quantization removed a single scale across the vector, and everything
here normalises. A bit table cannot even be read — it packs one bit per
component and bit[7] is a legal declaration, so the number of components is not
recoverable from the blob and there is nowhere else to read it from. The backend
declares can_read_vectors=False and says why.
What the shadow copy holds. rebasis shadows what the store returns, and a quantized store returns a value decoded from its stored code. The shadow is still bit-identical — to that decoded view. It is not a copy of the vectors your embedding model produced; those stopped being retrievable when the collection was built.
What rollback therefore restores. The state the migration replaced,
exactly: the vectors this collection read back the moment before migrate
started. It does not recover precision the collection had already spent.
It can stop the run. After every batch migrate re-reads a sample and
compares it to what it sent, to VERIFY_ATOL — the constant is in
src/rebasis/migrate/engine.py, and migrate --dry-run prints whatever it
currently is rather than a figure copied into prose. That check exists to catch
a store that accepts a write and does not keep it; a store that re-encodes on
write fails it for a different reason. Measured in
tests/integration/test_quantized_roundtrip.py, an 8-bit scalar-quantized FAISS
index deviates by more than that tolerance in both directions — so on a codec
that coarse the job stops on its first batch, with the shadow copy already
written and nothing lost.
What each backend reports¶
StoreCapabilities.quantized has three values, and the third is the point.
False is a promise that what you write is what you read back; a backend that
answered False without looking would be making a guarantee it could not keep.
So the default is None — not determinable — and that is what a third-party
store behind the LangChain or LlamaIndex bridge honestly is.
| Backend | Read from | Reports |
|---|---|---|
pgvector |
format_type on the column, from the catalogue |
vector → False; halfvec, bit, sparsevec → True |
faiss |
sa_code_size() against 4 × d, through the IndexIDMap2 wrapper |
True for PQ, scalar-quantized, LSH and friends; False for a flat index |
sqlite-vec |
vec_type() on a stored vector |
float32 → False; int8 and bit → True; empty table → None |
qdrant |
VectorParams.datatype in the collection config |
float16, uint8, turbo4 → True; otherwise False |
lancedb |
the Arrow element type of the vector column | narrower than 32 bits → True |
chroma |
nothing to read | False |
memory |
nothing to read | False |
Three of those answers are deliberately narrower than they first look, and all three are worth knowing if you go looking for a warning that does not appear.
pgvector's halfvec stops a migration rather than degrading one. Measured
(tests/integration/test_pgvector_types.py): a halfvec column rounds by more
than migrate's 1e-4 read-back tolerance, so a migration into one trips its own
verification on the first batch. The pre-flight plan says so beforehand rather
than letting it surface as a failed write. bit and sparsevec are refused
outright — they hold a code rather than a reconstruction of the vector that
produced it, and there is nothing for migrate to write back into.
format_type rather than information_schema.columns.data_type, because the
latter reports every extension type as USER-DEFINED and cannot tell vector
from halfvec — which is the entire question.
Qdrant's quantization_config is not this. Qdrant builds the quantized
codes beside the vectors rather than instead of them — which is what makes its
own rescoring possible — so a scalar- or binary-quantized Qdrant collection
still returns the original vector, and rollback on one is exact. Qdrant draws
the line itself: "datatypes are distinct from the quantization feature.
Quantization creates a separate quantized representation of vectors alongside
the original ones, while datatypes determine the representation of the original
vectors themselves." So it is datatype that rebasis reads.
LanceDB's IVF_PQ is not this either. The compressed copy lives in the
index's own columns and the vector column is untouched, which is why LanceDB
documents bypass_vector_index as a way to get ground-truth results. What does
change the round trip there is a vector column that is not float32 — uint8
columns are a supported way to store binary embeddings — and that is what is
read.
FAISS is the one where a quantized index used to pass silently. rebasis
already refuses a FAISS index it cannot reconstruct from, but that catches only
the ones where reconstruct raises, such as an IVF index with no direct map.
An IndexPQ or an IndexScalarQuantizer reconstructs happily and returns a
decode. Those are now declared rather than refused.
Reading it from a script¶
migrate has no --json. The finding is written into the audit trail with the
job, as store_quantized among the inputs of the migrate.job.started record:
rebasis audit export --out trail.jsonl
jq 'select(.action == "migrate.job.started") | .inputs.store_quantized' trail.jsonl
true, false and null are three different answers there, and null means
the store could not be asked — not that it keeps what it is given.
--no-keep-original¶
Disables the shadow copy. It saves disk and it makes the migration irreversible.
rebasis will not let that happen quietly: it prints a warning before starting and records the choice in the audit trail. If the result is not what you expected afterwards, the original vectors are gone.
Cleaning up¶
A dry run by default. A garbage collector that deletes without being asked is the exact class of data loss it exists to prevent.
Removing a shadow copy makes that job permanently irreversible, so it needs
--i-understand on top of --apply.
When a batch fails¶
The run does not stop. The records that could not be written are marked
FAILED with their error code, the ones that could are marked DONE, and the
next batch starts. A run therefore finishes with a count of failures rather than
at the first one, which is what makes a partial network outage recoverable in one
resume instead of many.
Two things happen before anything is given up on.
The batch is retried. StoreWriteFailed declares itself transient — a store
that refused a write because a node was rebalancing will usually take it a moment
later — so the write is attempted three times with exponential backoff and
jitter. Every attempt after the first is logged, so a run that took four extra
seconds says why.
Then the batch is split. If it still fails, it is halved and each half is written separately, recursively, so the records that were always fine get written and only the ones the store actually refuses end up in the failed list. A single oversized payload used to cost its two hundred and fifty-five neighbours a place in that list. Splitting is bounded at four levels: a store that is simply unreachable fails every half, and splitting all the way down would cost 511 writes to learn what the first one already said.
The halves are not retried. The batch's own three attempts already established that waiting does not help; what is left to find out is which record.
Look at what failed, fix the cause, then rebasis resume <job-id> — a resume
returns failed records to the queue.
The one thing that never happens is a partially-written record: the shadow is taken first, the write is one call, and the read-back verifies it.
A runbook¶
What to do when a migration goes wrong, in the order to do it. Every command here is read-only unless it says otherwise.
It stopped and you do not know why¶
state is what the engine last recorded. Six things stop a run, and status
distinguishes them:
state |
pause_requested |
What happened | What to do |
|---|---|---|---|
running |
false |
It is still going | Nothing |
running |
true |
A pause or a signal is in flight | Wait for the batch |
paused |
— | Asked to stop: pause, SIGTERM, low battery, memory pressure |
rebasis resume <job-id> |
paused with failed > 0 |
— | A batch failed and the job stopped | Read the error code first |
completed |
— | Nothing is wrong | — |
failed |
— | It could not continue | Read the error code |
If rebasis status shows nothing, you are pointed at the wrong state directory.
REBASIS_STATE_DIR, or --state-dir.
The process is gone and the lock is still held¶
The lock records the holder's PID, the operation and the start time, and rebasis
will tell you whether that process still exists. It will not break the lock for
you. Being wrong about a process being dead is how two writers end up in one
manifest — so if it says the holder is gone, remove
<state-dir>/rebasis.lock yourself, having checked.
Queries got worse and nothing else changed¶
The most likely cause is the one nothing else detects: the index is holding two embedding spaces at once, because a migration stopped part-way.
Measured, an ordinary bridged query against a half-migrated collection drops from a hit rate above 0.90 to below 0.65 — silently. Three ways out, in order of how much you have to change:
- Finish the job.
rebasis resume <job-id>. The condition ends when the queue empties. - Serve it correctly meanwhile.
rebasis.serve.MixedSpaceSearchrestores it to above 0.90 at every stage — see Searching one anyway. - Go back.
rebasis rollback <job-id>, below.
A batch failed¶
The error code says which of these it is, and every code is in the error reference.
| Code | Usually means | What to do |
|---|---|---|
RB-E6002 |
One or more records failed | The ids and their codes are in the queue; rebasis status counts them |
RB-E3004 |
The store rejected a write | Read the store's own logs. Everything already written stays written |
RB-E6004 |
Out of disk | The pre-flight sized the job; something else filled the disk since |
RB-E6005 |
The store took a write and did not keep it | Stop. This is the durability check failing on a fresh connection — do not resume until you know why |
RB-E3005 |
The adapter's width does not match the index's | The wrong adapter. Check rebasis eval <adapter> --verify before resuming |
RB-E4002 |
The adapter refuses this index | Fingerprint mismatch, or the wrong direction — migrate needs --direction old_to_new |
RB-E7003 |
The shadow copy is missing | Rollback cannot restore. Do not run --no-keep-original again |
RB-E7004 |
The state lock is held | Another writer, or a stale lock — see above |
Fix the cause, then rebasis resume <job-id>. Everything already written stays
written; everything not yet written stays queued.
You want it undone¶
Restores from the shadow copy, which is taken before each batch is overwritten. Two things to know before you rely on it:
- It restores the vectors, not the index structure. On a graph backend, run
--rebuild-indexafterwards if the health check reports a drop. - If the job ran with
--no-keep-originalthere is no shadow and no way back. That flag needs--i-understandfor this reason.
You are not sure the rollback worked¶
Read-only in every path. It opens the store, checks the SQLite file's integrity,
reads the recorded encoding profile, and says whether the collection holds two
embedding spaces. Run it first whenever anything is confusing — it is also the
thing to attach to a bug report, as --json.
None of the above¶
rebasis doctor --json and an issue. It carries no document text, no vectors and
no credentials by construction, so it is safe to paste.