发布于 2026年10月09日 · 我们于 2026年10月09日 确认该职位仍然有效
这是您的公司吗?₹ 600 – ₹ 1.500 (每个项目)
I have weekly supermarket flyers in PDF form—around 40 pages each—containing product images, names, normal price, offer price, pack size/quantity, validity dates of the promotion and any additional information. I need an offline-first API that takes an uploaded PDF, scans every page and returns those fields together with a confidence score for each value. Extracted date - Product names, normal price, offer price, pack size/quantity, validity dates of the promotion, any additional information, and confidence score. Core requirements • Language: feel free to build in Python or Node.js or Java, whichever lets you reach higher accuracy and faster throughput. • Absolutely no external calls: the full pipeline (OCR, layout analysis, model inference) must work without internet access once installed. • Paging & resume: the service should checkpoint page-level results to disk so that, in case of interruption, processing can restart from the last completed page. • Performance: processing time per page should stay reasonable; the entire 40-page sample flyer must finish in one run on commodity hardware. • Output: selectable JSON or CSV. Each record must include page number, product name, details, size/quantity, normal price, offer price, valid-from, valid-to, plus a confidence score for every field. • Stateless HTTP endpoints: standard POST for upload, GET for status/progress, GET for page-level results. No authentication layer is required. • Clear API documentation and a short setup guide so the system can be installed on an air-gapped server. Deliverables 1. Source code with install script (Python or Node.js). 2. Trained models or rule sets required for offline extraction. 3. API documentation (OpenAPI / Swagger preferred). 4. README covering local deployment, configuration of checkpoints, and example curl calls. 5. Final run on the attached sample PDF demonstrating successful extraction and resume capability. Acceptance will be based on extraction accuracy, processing speed, ability to restart mid-run without data loss, and conformity to the documented endpoints.