Published on Sep 25, 2026 · We confirmed on Sep 27, 2026 that it's still live
Is this your business?US$ 30 – US$ 250 per project
I need a robust, from-scratch search engine that can crawl and index the full spectrum of online content—standard web pages, document formats such as PDF and DOC(X), and rich media files (images, audio, video). Every language matters; the pipeline must recognise, store, and make searchable content written in anything from English and Spanish to lesser-used scripts. Core expectations • A distributed crawler that respects robots.txt while gathering pages, docs, and media. • A scalable indexing layer (Lucene, Elastic, Solr, or comparable) tuned for multilingual tokenisation and ranking. • Relevance algorithms that handle mixed content types and surface quality results quickly. • An intuitive search interface (web-based) with filters for file type, language, and date. • Clear build/run documentation plus deployment scripts so I can spin the stack up on my own servers later. Completion is accepted when the system reliably: 1. Crawls a test set of sites in multiple languages without errors. 2. Returns accurate, ranked results across web, document, and media queries. 3. Survives load testing at scale and passes code review for clarity and security.
Create a free account to see the full job and apply.