We’d been paying for a separate vector store to sit next to a database that already held every document we were embedding. When the DiskANN-based vector index and VECTOR_SEARCH graduated to GA in Azure SQL Database on 28 September, the case to collapse those two systems into one basically wrote itself. Same durability guarantees, one backup story, one network boundary, one thing to patch. And GA is not cosmetic here — it means a production SLA and Microsoft support now stand behind the feature, not just that the syntax parses. So I moved it. Here’s what actually happened.
The dimension wall, about ten minutes in
The vector data type has a dimension limit, and our embeddings blew past it — 3,072 dimensions, the default for text-embedding-3-large. The ALTER TABLE to add the column failed before I’d written a single line of search code.
This is the one nobody warns you about because everyone demos with text-embedding-3-small, which is 1,536 and fits fine. The fix is not to re-pick your model in a panic. The v3 embedding APIs take a dimensions parameter — you ask for a smaller vector and get a shortened, still-usable one. But that means re-embedding the entire corpus. Budget for that. It’s a batch job and a bill, not a config flag.
The query that worked in dev and died in prod
The vector type and VECTOR_DISTANCE already existed before this release — what’s new in GA is the index. So I already had exact nearest-neighbour queries running:
That is a full scan. Every row, every query. At 50k chunks in dev it returned instantly and I moved on, pleased with myself. At ~4M chunks in production it fell over. Exact KNN doesn’t scale, and that limitation is the entire reason the index is the news here, not the distance function.
Approximate nearest neighbour (ANN) is the trade. DiskANN doesn’t guarantee it found the ten closest vectors — it finds ten very likely to be closest, and it does it without touching every row. You give up a sliver of recall for orders-of-magnitude less work. For RAG that’s the right trade; you’re feeding a language model context, not settling a lawsuit. Create the index, then query through VECTOR_SEARCH:
Match the metric on the index to the one in the query, and match both to how your model was trained — cosine for the OpenAI embeddings, which is what most people are running. Get that wrong and you get plausible-looking garbage, which is worse than an error.
The freshness surprise, the next morning
New documents ingested overnight weren’t showing up in results. The ANN index is not the live column — it’s a structure built over it, and newly written vectors aren’t instantly reflected the way a b-tree update would be. If your pipeline writes embeddings continuously and expects them searchable in the same breath, test that assumption explicitly before you promise it to anyone. We moved to a batched ingest window and stopped fighting it.
What I’d tell you before you start
Separate two questions: which tiers can run this at all, and which run it well. Treat those differently. The VECTOR_SEARCH surface itself is the same T-SQL regardless of where you sit, but DiskANN is memory-hungry, so a database with real RAM behind it — vCore or Hyperscale — is where you want to be for a production corpus, not the bottom DTU tiers where you’ll discover the constraint the hard way. Confirm the current tier and region support against Microsoft’s docs for your subscription before you plan the cutover; GA does not mean everywhere on day one.
And keep your expectations honest: this does not turn Azure SQL into Pinecone. If vector search is your product and you’re running billions of vectors with aggressive recall tuning, a dedicated engine still earns its keep.
But that’s not most of us. Most of us have a few million chunks living next to relational data we’re already filtering and joining on — tenant ID, document status, ACLs. Doing that filtering and the similarity search in one SELECT, under one SLA, with no second system to sync, is the actual win. I deleted a vector store this week. The RAG results got no worse. The architecture diagram got a lot shorter.
