Hi all,
I run inetgeek.com, a small comparison site for hosting and infrastructure
providers. I’ve started using the Technology Report API to show per-category
adoption bars — for example the three PaaS providers I track, sized by origins
detected in the latest crawl, linking back to the relevant tech report page.
Two things I couldn’t resolve from the docs, and I’d rather ask than assume:
-
Is there a licence for the Technology Report / HTTP Archive dataset itself?
I checked the FAQ, the About page, har.fyi and the GitHub org. I can see
httparchive.org and data-pipeline are Apache-2.0 and the Wappalyzer fork is
GPL-3.0, but those cover code rather than the output numbers. Is the data
CC0, CC-BY, or something else? -
Is there a preferred attribution string? Right now each chart reads
“Source: HTTP Archive” and links to
HTTP Archive: Tech Report . Happy to change the wording
or the link target if you’d rather it pointed somewhere specific.
For what it’s worth, usage is tiny: one build-time request per month, and the
numbers are cached in the repo rather than fetched per page view.
One small note in case it’s useful to the team: a few technologies that exist in
the catalogue are easy to mistake for a different vendor of the same name —
“Doppler” is fromdoppler.com (email marketing) rather than doppler.com (the
secrets manager), and “Neon CRM” is Neon One rather than neon.com (Postgres).
Not a bug, just a trap for anyone matching vendor names automatically.
Thanks for publishing this openly — it’s the only free source I found that
counts the same kind of thing as the commercial crawlers.