pgvector¶
The store most teams already run, and the only one of the six where the migration's weakest guarantee is the database's rather than rebasis'.
The driver is pg8000, not psycopg, and the reason is a licence rather than
a feature. psycopg is the obvious choice and this backend was written against
it; it is LGPL-3.0-only, and rebasis is Apache-2.0 and meant to sit inside
products a copyleft dependency would exclude it from — so ci.yml's dependency
review denies that family, and caught it on the first pull request. pg8000 is
BSD-3-Clause and pure Python. It has no named-cursor API and no
connection.transaction(), so the two places that need them issue
DECLARE/FETCH and BEGIN/COMMIT themselves.
The URI¶
pgvector://user:password@host:5432/dbname#public.documents
pgvector://user@host/dbname#documents # schema defaults to public
The fragment is the table, optionally schema-qualified. The database is the
path. Credentials are parsed out of the URI and never reach a log line, an
audit record or an error message — rebasis doctor --json, which the README
tells people to attach to a bug report, prints pgvector://<credentials>@host.
Which column is which¶
This is the part that needs a decision, and it is pgvector-specific.
A Chroma collection has a shape. A Postgres table is whatever its owner made, so rebasis has to be told — or find the conventional name and say which it found:
| tried, in order | |
|---|---|
| vector | embedding, vector, embeddings, vec |
| id | id, doc_id, chunk_id, key, uuid |
| text | text, content, document, chunk, body |
When your schema uses something else, name it:
Nothing is inferred from a column merely being the only one of its type. A
table with two vector columns would otherwise have migrate rewrite whichever
one the catalogue happened to return first. If no conventional name is there,
the error lists the table's own columns with their types and names the option
that supplies the missing one.
A missing text column is not an error — it is can_read_text: false, and it
means probe and the bridge still work while Cascade and any fit that needs
to re-embed do not. Check what was resolved before relying on it:
The column type decides almost everything¶
pgvector has four vector types and they are not interchangeable. rebasis reads
the declared type from the catalogue — through format_type, because
information_schema reports every extension type as USER-DEFINED and cannot
tell vector from halfvec — and behaves differently per type.
| type | what it holds | read | write | quantized |
|---|---|---|---|---|
vector |
float32 | yes | yes | false |
halfvec |
float16 | yes | yes | true |
bit |
a binary code | no | no | true |
sparsevec |
an index/value map | no | no | true |
Measured, not assumed (tests/integration/test_pgvector_types.py):
- A
vectorcolumn returns exactly what was written. That is the promiserollbackrests on. - A
halfveccolumn rounds, and by more thanmigrate's read-back tolerance of 1e-4. So a migration into ahalfveccolumn would stop on its own first batch — the pre-flight plan says so beforehand rather than letting it surface as a failed write. Read If your index is stored quantized. bitandsparsevechold a code rather than a reconstruction of the vector that produced it, and pgvector scores them with Hamming, Jaccard or its own sparse operators rather than the cosine distance rebasis speaks. Both are refused with the reason, at the moment they are asked.
The precedent for reading the type rather than reasoning about it is
sqlite-vec, where an int8 column was once read as if it held float32 and
every number derived from the reported dimension was a quarter of the truth.
Nothing raised.
What the transaction takes over¶
migrate's durability chain is four mechanisms: a shadow copy before every
batch, a read-back after, a fresh-connection check when the queue empties, and
rollback. All correct, and all durability rebuilt above the storage engine.
On pgvector, one batch is one transaction. It lands whole or it does not
exist, so the half-written batch stops being a state anything has to detect or
undo. migrate's pre-flight plan names this, because knowing which layer holds
which guarantee is the difference between trusting a tool and hoping.
The shadow copy stays, and that is deliberate. A transaction rolls back one
batch; rebasis rollback <job-id> rolls back a finished job three days later.
Different scopes, and the second is the one a user asks for.
The connection is opened in autocommit, which is the other half of the same
decision. DB-API's default opens a transaction on the first statement and holds
it until something commits — so a probe against a live database would sit
idle in transaction for the length of the run, pinning the vacuum horizon of a
production table for a read that needed no transaction at all. The two places
that genuinely need one open it explicitly.
Rebuilding the index¶
pgvector is the second backend that can rebuild its own search structure, and the first where the two index types answer differently.
rebasis migrate --adapter adapters/forward.rbs \
--store "pgvector://user@host/db#public.documents" --rebuild-index
REINDEX INDEX CONCURRENTLY, not a plain REINDEX. The plain form takes an
ACCESS EXCLUSIVE lock and stops every read of the table for its duration, which
is not a thing a tool should do to somebody's production index. The concurrent
form builds the replacement beside the original and swaps, at the cost of more
disk and a longer run. If it fails part-way it leaves an invalid index behind;
\d on the table names it and DROP INDEX removes it. The vectors are
untouched either way.
Which index you built decides whether this is insurance or part of the job.
Measured on 100,000 records with the orthogonal map auto picks
(the numbers):
| before | after | after REINDEX CONCURRENTLY |
|
|---|---|---|---|
| HNSW | 0.970 | 0.913 | 0.956 |
| IVFFlat | 0.853 | 0.308 | 0.838 |
(Re-measured under the shipped pg8000 driver at 0.896 → 0.322 → 0.875 and
0.964 → 0.893 → 0.960; the table above is the psycopg run these were checked
against, and the difference is the repeat spread of a non-deterministic index
build.)
On IVFFlat a migration costs two thirds of the index's recall, because its list centroids were computed once from a distribution the migration rotated. Every vector is correct and in the wrong list. The rebuild recovers essentially all of it. On HNSW the loss is 6 points and the rebuild recovers most.
IVFFlat also migrates two to four times faster, because maintaining a list assignment on write is cheaper than maintaining a graph. That is the trade: a slower migration that degrades a little, or a faster one that degrades a lot and must be reindexed.
A table with no index on the vector column has nothing to rebuild, and that is not a failure — an exact scan cannot lose recall.
Reads stream¶
iter_records opens a server-side cursor and fetches batch_size rows at a
time. A client-side cursor reads the whole result set into the
client before yielding the first row, which is exactly the O(N × d) peak the
store contract forbids. A five-million-row table costs the same resident memory
as a fifty-thousand-row one.
What pgvector supports¶
| Capability | Supported |
|---|---|
| Read vectors | yes, on vector and halfvec |
| Read text | if a text column was found or named |
| Upsert vectors | yes, on vector and halfvec |
| Metadata filter | yes — a SQL WHERE, the richest of the six |
| Dimension locked | yes |
| In-place update | yes |
| Rebuild index | yes |
The locked dimension is vector(n)'s own type modifier. Changing it is DDL, and
DDL is the thing rebasis does not do: migrate changes vectors, not schemas. A
move from vector(1024) to vector(256), or from vector to halfvec, is
yours to perform — and probe --truncate is what tells
you whether it is worth performing.
At a million rows¶
ROADMAP.md's largest stated gap is one sentence: everything is tested on
hundreds of records, not millions. pgvector is the first backend where a table
of that size can be stood up in minutes on a machine the project already has,
so it was — spikes/pgvector_scale.py, 1,000,000 rows at 384 dimensions.
| table | migrated | seconds | records/s | peak traced memory |
|---|---|---|---|---|
| 20,000 | 5,000 | 10.5 | 478 | 7.3 MB |
| 1,000,000 | 5,000 | 14.1 | 354 | 7.3 MB |
| 1,000,000 | 100,000 | 342.6 | 292 | 17.4 MB |
The middle row is the finding. Migrating the same 5,000 records against a
table fifty times larger peaks at exactly the same 7.3 MB. Peak memory is a
function of how many records are enqueued, not of how many are in the table
— which is the O(batch × d) streaming contract, measured at a million rows
rather than argued from a docstring. The third row's extra 10 MB is 95,000 more
queued ids, about 105 bytes each, which is a Python string in a list.
The shadow scaled with it and stayed correct. 100,000 records, 153.6 MB on
disk, every migrated record covered, and a sample read back cleanly. That is
rollback's promise at a size nobody had asked it about.
Two things this does not establish, and they are the half of the ROADMAP's sentence that stays open:
- The table had no index on the vector column. These figures are the write
path alone. Index maintenance on
UPDATEis real and measured separately — what a migration does to the index puts HNSW at two to four times IVFFlat's cost at 100,000 records — so a migration against an indexed million-row table will be slower than 292 records per second. - It is still not somebody's index. A synthetic million rows on a machine
running three other jobs is not a production database, and nobody has yet
pointed
migrateat an index they could not rebuild. That gap needs somebody else's data and a backup, and no amount of measuring here closes it.
One practical finding from standing the table up: building an IVFFlat index
over a million 384-dimensional vectors asks for 76 MB and PostgreSQL's default
maintenance_work_mem is 64, so the build fails — "memory required is 76 MB,
maintenance_work_mem is 64 MB" — after the whole table has loaded. rebasis
never creates an index, so this is not its failure; it is what anybody creating
a vector index at that size meets. SET maintenance_work_mem = '1GB' for the
session that builds it.
Scope, stated plainly¶
One table, named columns. A joined schema, a partitioned table, text in a second table: all refused explicitly rather than half-supported silently. The narrowing is the point — a backend that half-works on a schema shape is worse than one that says which shape it works on.
This is a real database under a live application, which the other backends
mostly are not. --dry-run prints the plan without writing anything, and
--limit migrates a slice. Both are worth using here even if you skip them
elsewhere.