TL;DR Start with a strict embedding contract Jobs turn a database snapshot into a vector corpus Buckets are the connective tissue Inference Endpoints put semantic search on the request path Hybrid retrieval is stronger than either branch alone One Endpoint, two update paths Related papers become almost free online What we learned 1. Separate throughput work from latency-sensitive work 2. Make storage the explicit contract between compute and production 3. Pin more than the model name 4. Design for cold starts 5. Smaller vectors can be a systems feature 6. Activation should be boring 3 months ago, we started a revival of Papers with Code (see also the announcement tweet). Its goal is to make open AI research accessible and digestible, so that people can easily find the artifacts related to a paper, find state-of-the-art (SOTA) across the various domains of AI, share interesting research and build on top of each other's work. In other words, its goal is to power the wave of research that leads to the next Transformer.