pgvector vs Pinecone for Node.js Embeddings
pgvector vs Pinecone comes up on basically every Node.js project I've touched that needed semantic search or retrieval-augmented generation in the last year, and the honest answer is that most teams pick the wrong one for their actual traffic pattern. I built a document-search feature for a client's internal knowledge base last quarter and we went back and forth on this exact decision twice before landing on something that actually fit the project, not the one with the flashiest landing page.
There are two ways people usually frame this choice, and both of them skip the part that actually matters: it's not "which is better," it's "which one matches the size and shape of your data." Here's how I'd actually walk through it if you're deciding right now.
What each one actually is
pgvector is a Postgres extension. If you already run Postgres — which, if you're building a Node.js backend with Prisma, Knex, or raw pg, you probably do — it adds a vector column type and similarity search operators (<-> for L2 distance, <=> for cosine distance) directly to your existing database. No new service, no new API key, no new bill.
Pinecone is a dedicated, managed vector database. You send it vectors over its API, it indexes them with algorithms built specifically for approximate nearest-neighbor search at scale, and it hands results back fast even across tens of millions of vectors. It's a separate system from your Postgres database, with its own pricing, its own SDK, and its own operational surface.
| pgvector | Pinecone | |
|---|---|---|
| Infrastructure | Extension on your existing Postgres | Separate managed service |
| Setup cost | Near zero if you already run Postgres | New account, API key, SDK integration |
| Query latency at <1M vectors | Fast, especially with an HNSW index | Fast |
| Query latency at 10M+ vectors | Degrades unless carefully tuned | Built for this scale |
| Joins with relational data | Native — one query, one database | Requires a second round trip to Postgres |
| Operational overhead | None beyond your existing Postgres ops | Monitoring a second vendor, rate limits, billing |
| Cost at small-to-medium scale | Effectively free (just storage) | Starts at $0 free tier, climbs with vector count |
Where pgvector wins
If your vector count is under a few million and you already have a Postgres database holding the rest of your app's data, pgvector is almost always the right call, and I'd actively push back on anyone reaching for Pinecone by default here. The real win isn't raw query speed — it's that you can join the vector search against your relational data in one query, in one transaction, with no second network hop and no eventual-consistency gap between "the document was saved" and "the document is searchable."
const { Pool } = require('pg');
const pool = new Pool({ connectionString: process.env.DATABASE_URL });
async function searchSimilarDocuments(queryEmbedding, userId, limit = 5) {
// queryEmbedding is a 1536-length float array from your embedding model
const result = await pool.query(
`SELECT id, title, content, 1 - (embedding <=> $1::vector) AS similarity
FROM documents
WHERE user_id = $2
ORDER BY embedding <=> $1::vector
LIMIT $3`,
[JSON.stringify(queryEmbedding), userId, limit]
);
return result.rows;
}
Notice that WHERE user_id = $2 sitting right next to the vector search. That's the part Pinecone can't do as cleanly — you'd either need to pass user ID as metadata and filter server-side with its own query syntax, or fetch candidate IDs first and filter in a second step. For most SaaS apps doing per-tenant search, that single detail is enough to settle the decision on its own.
For index performance at this scale, add an HNSW index once you're past a few thousand rows:
CREATE INDEX ON documents USING hnsw (embedding vector_cosine_ops);
Where Pinecone actually earns its cost
I'd stop recommending pgvector the moment you're genuinely looking at tens of millions of vectors with high query-per-second requirements, or when vector search is the entire product, not a feature bolted onto an existing app. Pinecone's indexing algorithms are purpose-built for that scale in a way that pgvector's HNSW implementation isn't tuned for out of the box — you'd need to seriously invest in index parameter tuning, sharding strategy, and read replicas to get pgvector performing well at that volume, and at that point you're half-building what Pinecone already gives you.
The other legitimate case: you don't want vector search coupled to your primary database's resource budget at all. If a surge in embedding queries competing for the same CPU and memory as your transactional workload is a real operational risk for you, running vector search on a separate system is a reasonable tradeoff even below the scale where it's strictly necessary.
My actual recommendation
Default to pgvector unless you have a specific, measured reason not to. Most teams evaluating this are nowhere near the scale where Pinecone's advantages show up, and the cost of adding a second managed service — a new bill, a new failure mode, a new thing to monitor at 2am — is real and usually underestimated at the planning stage. Start with pgvector, instrument your actual query latency and vector count in production, and migrate to Pinecone (or a comparable dedicated vector store) only when you have real numbers showing you've outgrown it, not because a blog post said vector databases are the modern way to do this. I've seen more projects over-engineer this decision upfront than actually hit the scale where it mattered.
Frequently Asked Questions
Can I switch from pgvector to Pinecone later without rewriting everything? Mostly yes, if you've kept your embedding logic separate from your storage logic. The embeddings themselves don't change — you're just moving where they're stored and queried from. Budget time for rewriting your search query layer, since the filtering syntax and client libraries are completely different.
Does pgvector support the same similarity metrics as Pinecone? Yes — L2 distance, cosine distance, and inner product are all supported by both, so metric choice isn't a factor in picking between them.
What embedding dimension should I use with pgvector?
Whatever your embedding model outputs — commonly 1536 for OpenAI's text-embedding-3-small. pgvector handles up to 2000 dimensions efficiently with HNSW; beyond that you'll want to check the current version's limits since this has changed across releases.
Is Supabase's pgvector support the same as self-hosted pgvector? Functionally yes — Supabase just runs standard Postgres with the extension enabled, so everything here applies directly if that's your hosting setup.
Related posts
Node.js RAG Pipeline: pgvector + OpenAI
Build a Node.js RAG pipeline with pgvector and OpenAI embeddings for grounded, source-backed AI answers without running a separate vector database.
Caching OpenAI API Responses in Node.js
Caching OpenAI API responses in Node.js with Redis: exact-match vs semantic caching, the real tradeoffs, and code that actually cuts your costs.
Stream OpenAI API Responses in Node.js with SSE
Learn how to stream OpenAI API responses in Node.js using Server-Sent Events, with real Express code and the mistakes that break streaming in production.
Fix OpenAI API Rate Limit Errors in Node.js
Get 429 rate_limit_exceeded errors from the OpenAI API in Node.js? Here is the real fix: exponential backoff, retry-after handling, and concurrency limits.
0 Comments
No comments yet — be the first to share your thoughts.