Papers with Code search relies on a pipeline of Hugging Face managed services to turn research papers into searchable vectors. The system uses Inference Endpoints to serve the embedding model, which converts queries and paper text into numerical representations. This keeps the live search responsive without requiring dedicated GPU infrastructure.
Behind the scenes, Jobs handle the batch work of embedding newly ingested papers on a schedule. This separates real-time inference from background updates, so the index stays fresh without slowing down user requests. Buckets provide object storage for the vector index and associated metadata, making the data available to the search service at query time.
The architecture shows how combining three managed components—endpoints, jobs, and buckets—can support a full search feature. By splitting compute into synchronous and asynchronous paths and using managed storage, the team avoids custom infrastructure while scaling with the corpus of papers. The post is a concrete example of using Hugging Face's platform beyond model hosting.