نُشرت في 29 سبتمبر 2026 · تحققنا في 29 سبتمبر 2026 من أنها ما زالت متاحة
هل هذه شركتك؟US$ 250 – US$ 750 / لكل مشروع
Title: Python crawler for Norwegian municipal case archive (Saksinnsyn) — Docker, Postgres, resume, provenance Brief: I need a senior Python developer to build a polite, resumable crawler for Oslo municipality's public building-case archive (innsyn.pbe.oslo.kommune.no/saksinnsyn). Plain HTTP (requests/httpx + lxml), no browser automation. I have a working single-file prototype (v0.29) that covers discovery, case pages, journalpost lists, attachment download, manifest and block detection; you may reuse or rewrite it, your choice, the price is fixed either way. Scope Discover all byggesaker 2019–present via the site's search (I'll give you the working query patterns and the case-number rules, which differ before/after 2025 and must be config per source and year, not hardcoded). Per case: save the raw case page HTML plus parsed fields (saksnummer as shown + normalized, title, sakstype, status, dates, address, gnr/bnr, unit). Per case: full journalpost list, raw HTML + parsed rows (number, date, title, inn/ut/intern, sender/recipient). Per journalpost: every attachment (file id, filename, file). Items that are login-gated or withheld are recorded with status and reason, not fetched. Linked cases (kobling): record the link and fetch the linked case. Every download logged: URL, time, HTTP status, size, sha256, run id. Output: raw files on disk (verbatim, never rewritten) plus a Postgres table case → journalpost → document → file → status (fetched / gated / withheld / missing / blocked). Idempotent and resumable: re-running never duplicates; stop/start any time. Multi-municipality by design: every record carries kommunenr and source; Oslo fetching is one module behind a common data model. Nothing downstream may know about Oslo's pages. Delivered as Docker Compose (crawler + Postgres) running on my Windows PC, CLI-driven. No web UI. Site behaviour you must handle (this is the hard part) The server intermittently blocks file downloads (showfile.asp) with a captcha page or a redirect to ID-porten login. This is server-side and not IP-based; it also hits a normal browser. Required behaviour: detect both responses, mark the file blocked with reason/time/response hash, keep fetching metadata pages (which stay open), probe at most once per hour, resume automatically when the gate reopens. Fixed User-Agent string that I supply, single session, rate limit configurable (default 1 request / 3 s), daily file budget configurable. No captcha solving, no ID-porten automation, no proxy/IP/session/User-Agent rotation. Bids that propose any of these will not be considered. I'll test this with a mock server returning the block page. Acceptance End-to-end run on 200 cases I specify produces the status table and matching files. Full 2024 run matches sha256 of ~3,200 files I already hold. Block simulation: crawler stops, logs, resumes, no duplicates. Terms Fixed price, milestones 30% start / 40% at 200-case run / 30% at 2024 verified. Code in my GitHub repo from day one, IP assigned to me. Timeline 3 weeks. You should have: 5+ years Python, prior crawling of government or public-sector archives, Postgres, Docker. Say in your bid what you'd do when the crawler gets the captcha page — I read that line first.
أنشئ حسابًا مجانيًا لعرض الوظيفة كاملة والتقديم عليها.