jobhunt-platform/packages/connectors/README.md
hermes 0e13ee5d51 feat(W1): packages/connectors - AF API, generic URL, dedupe, normalizer
- Connector protocol with fetch(query: SearchQuery) -> list[RawPosting]
- ArbetsformedlingenConnector: official Platsbanken API, field mapping
  (headline->title, employer.name->company, webpage_url->url,
  description.text->description, id->external_id), region mapping,
  polite User-Agent
- GenericUrlConnector: readability extraction (prefers main/article,
  strips nav/footer/script/style), title/company guess, Cloudflare
  challenge detection -> UnsupportedSite
- dedupe helper: (source, url) + external_id keying
- normalizer: RawPosting -> JobPosting (job_posting table shape)
- 37 tests, all passing, recorded fixtures (no live network)
- README with usage and mock examples
2026-07-30 18:23:24 +00:00

3.9 KiB

connectors

Job source adapters that fetch postings from external sources and normalize them to a common JobPosting shape matching the job_posting database table.

Connectors

Arbetsformedlingen (Platsbanken)

Uses the official Swedish Public Employment Service API. Free, no API key required.

  • Endpoint: https://jobsearch.api.jobtechdev.se/search
  • Method: GET, header Accept: application/json
  • Field mapping: headline -> title, employer.name -> company, webpage_url -> url, description.text -> description, id -> external_id
  • Region mapping: common names (e.g. "malmo", "skane") mapped to API filter values
  • Polite User-Agent header

Generic URL

Fetches a single job posting URL and extracts clean text with simple readability extraction:

  • Prefers <main> or <article> containers
  • Strips nav, footer, script, style, aside, header, noscript tags
  • Extracts title from <h1> (falls back to <title>)
  • Attempts company name from og:site_name meta tag, then text heuristics
  • Does NOT follow Cloudflare challenge pages -- raises UnsupportedSite

Usage

Arbetsformedlingen

from connectors import ArbetsformedlingenConnector, SearchQuery

connector = ArbetsformedlingenConnector()
query = SearchQuery(query="python developer", region="malmo", limit=20)
raw_postings = connector.fetch(query)

# Normalize to JobPosting shape
from connectors import normalize
job_postings = [normalize(p) for p in raw_postings]

Generic URL

from connectors import GenericUrlConnector, SearchQuery

connector = GenericUrlConnector()
query = SearchQuery(query="https://example.com/jobs/123")
raw_postings = connector.fetch(query)
# Returns a single-element list

Deduplication

from connectors import dedupe

# Remove duplicates by (source, url) and (source, external_id)
unique_postings = dedupe(all_raw_postings)

Normalization

from connectors import normalize, RawPosting

raw = RawPosting(
    source="manual",
    external_id=None,
    url="https://example.com/job",
    company="Corp",
    title="Developer",
    description="A great job.",
)
job = normalize(raw)
# job.source, job.url, job.company, job.title, job.description, ...

Error handling

from connectors import UnsupportedSite

try:
    connector.fetch(SearchQuery(query="https://cloudflare-protected.com/job/1"))
except UnsupportedSite as e:
    print(f"Cannot scrape {e.url}: {e.reason}")

Data models

SearchQuery

Field Type Description
query str Search keywords or URL
region str | None Optional region filter
limit int Max results (default 20)

RawPosting

Field Type Description
source str Source identifier
external_id str | None Source-native ID
url str Posting URL
company str Company name
title str Job title
location str | None Job location
description str Job description text
raw dict Original source payload

JobPosting

Same fields as RawPosting. Matches the job_posting table shape: source, external_id, url, company, title, location, description, raw.

Development

cd packages/connectors
uv venv
. .venv/bin/activate
uv pip install -e ".[dev]"
pytest -q

Tests use recorded JSON/HTML fixtures (no live network calls). Fixtures live in tests/fixtures/.