opensourceprojects.dev

A broadsheet for software that doesn't ask for your email

Similarity search for billions of vectors, without keeping them all in RAM
GitHub RepoImpressions3

Project Description

View on GitHub

Similarity Search at Billions of Vectors Without Keeping Them All in RAM

You've got a pile of vectors and a query, and you need the nearest neighbors. Easy enough at a few million. But what happens when you're staring down billions of embeddings and the memory budget for your server is a hard ceiling? Faiss is one answer to that problem, and it's been the workhorse behind a lot of vector search systems for years.

What It Does

Faiss is a library for efficient similarity search and clustering of dense vectors. It assumes your instances are represented as vectors, identified by integers, and compared using L2 (Euclidean) distance or dot products. Vectors that are similar to a query are those with the lowest L2 distance or the highest dot product. Cosine similarity works too, since that's just a dot product on normalized vectors.

The library is written in C++ with complete Python and numpy wrappers, and some of the most useful algorithms are implemented on the GPU. It's developed primarily at Meta's Fundamental AI Research group. The core abstraction is an index type that stores a set of vectors and exposes a search function. Some index types are simple baselines like exact search; most of the interesting ones represent trade-offs across search time, search quality, memory per vector, training time, add time, and whether they need external data for unsupervised training.

The memory story is the part worth paying attention to. Methods based on binary vectors and compact quantization codes use only a compressed representation of the vectors and don't require keeping the originals around. That comes at the cost of less precise search, but it's what lets these methods scale to billions of vectors in main memory on a single server. Other methods, like HNSW and NSG, take a different approach: they add an indexing structure on top of the raw vectors to speed up search.

Why It's Cool

  • You don't have to choose between precision and scale. The compressed-representation methods trade accuracy for the ability to fit billions of vectors in RAM. The graph-based methods (HNSW, NSG) keep the raw vectors and layer an index on top for speed. Different problems, different tools, one library.

  • The GPU path is a drop-in replacement. If you've got GPUs, the GPU indexes swap in for CPU ones almost trivially — replace IndexFlatL2 with GpuIndexFlatL2 and the copies to and from GPU memory are handled for you. You'll get better performance if both input and output stay resident on the GPU, but the ergonomics of getting started are about as low-friction as it gets. Single and multi-GPU are both supported.

  • The dependency surface is small. The library is mostly C++, and the only hard dependency is a BLAS implementation. GPU support via CUDA or AMD ROCm is optional. The Python interface is optional. The NVIDIA cuVS backend implementations can be enabled optionally too. That's a lot of "optional" for a library this capable, which means you can pull in exactly the piece you need.

  • It's built for people who tune. The README is explicit that Faiss includes supporting code for evaluation and parameter tuning. This isn't a library that hides the knobs from you — the trade-offs are the point, and the tooling acknowledges that.

  • The documentation is actually organized. Full docs live on the wiki, with a tutorial, an FAQ, and a troubleshooting section. There's doxygen documentation at faiss.ai for per-class information pulled from code comments. For a library with this many index types and parameters, having a troubleshooting page is not a small thing.

How to Try It

The quickest route is the precompiled Anaconda packages:

  1. For CPU-only work: install faiss-cpu.
  2. For GPU work: install faiss-gpu.
  3. If you want the cuVS backend: install faiss-gpu-cuvs.

If you'd rather build from source, Faiss compiles with cmake. See INSTALL.md for the details.

Once it's installed, the natural starting point is the getting-started tutorial on the wiki. From there, the FAQ and troubleshooting section will cover most of the sharp edges you'll hit when you start picking index types and tuning parameters.

The repository is at github.com/facebookresearch/faiss. If you want per-class API details, the doxygen docs are at faiss.ai.

Final Thoughts

Faiss isn't trying to be a turnkey vector database — it's a library, and it expects you to understand the trade-offs you're making. That's the honest framing. If you need to search billions of vectors and you're willing to spend time on index selection and parameter tuning, it's a mature, well-documented option with a small dependency footprint and a GPU path that doesn't require rewriting your code. If you want something that just works out of the box with no tuning, this probably isn't it. For everyone in between, the wiki is the place to start.


Follow @githubprojects for more developer tools and open source projects.

Back to Projects
Project ID: 73a22404-7920-4aef-a272-991457c8f230Last updated: September 29, 2026 at 11:11 AM