# Best way to extract gtm.js response bodies at scale, BigQuery cost vs raw HAR access?

**URL:** <https://discuss.httparchive.org/t/best-way-to-extract-gtm-js-response-bodies-at-scale-bigquery-cost-vs-raw-har-access/3104>\
**Category:** Uncategorized\
**Created:** [July 7, 2026, 3:31pm UTC](https://discuss.httparchive.org/t/best-way-to-extract-gtm-js-response-bodies-at-scale-bigquery-cost-vs-raw-har-access/3104 "2026-07-07T15:31:26Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![dev.dandp](https://avatars.discourse-cdn.com/v4/letter/d/5f9b8f/32.png) [@dev.dandp](https://discuss.httparchive.org/u/dev.dandp)\
**Post date:** [July 7, 2026, 3:31pm UTC](https://discuss.httparchive.org/t/best-way-to-extract-gtm-js-response-bodies-at-scale-bigquery-cost-vs-raw-har-access/3104/1 "2026-07-07T15:31:26Z")

</div>

I’m doing research that requires collecting the **response bodies of `gtm.js` (Google Tag Manager) resources** across crawled sites, ideally over multiple monthly crawls.

My filter is essentially:

```sql
SELECT date, page, url, response_body
FROM `httparchive.crawl.requests`
WHERE client = 'desktop'
  AND is_root_page = TRUE
  AND type = 'script'
  AND CONTAINS_SUBSTR(url, 'gtm.js')
  AND response_body IS NOT NULL

```

The challenge is cost. Because `url` isn’t a clustering column, the `gtm.js` filter can’t prune, so extracting bodies means scanning the (very large) `response_body` column across each crawl. Over a wide date range this becomes expensive.

I’ve been trying to find a cheaper route and have a couple of questions:

1. **Is there a recommended pattern for extracting a narrow slice of `response_body` across many crawls cost-effectively?** (e.g. clustering tips I might be missing, a sampled/summary table, or a materialization approach the community uses.)

2. **Raw HAR access:** I understand the raw crawl data lives in `gs://httparchive` (`crawls/` and `crawls_manifest/`). I tried reading it with a billing project attached, but the bucket returns 403 on both `objects.list` and `objects.get`. Is there a process to request read access to the raw crawl data for? If so, what’s the right way to apply, and are there expectations around cost (requester-pays) or usage?

Thanks very much for maintaining such a valuable dataset, any guidance on the most cost-appropriate approach would be really appreciated.

---

<div class="post-metadata">

**Author:** ![max\_ostapenko](https://yyz1.discourse-cdn.com/flex035/user_avatar/discuss.httparchive.org/max_ostapenko/32/1515_2.png) [@max\_ostapenko](https://discuss.httparchive.org/u/max_ostapenko)\
**Post date:** [July 10, 2026, 11:13pm UTC](https://discuss.httparchive.org/t/best-way-to-extract-gtm-js-response-bodies-at-scale-bigquery-cost-vs-raw-har-access/3104/2 "2026-07-10T23:13:13Z")

</div>

Such analysis over a single monthly crawl costs around $5 - $10.

The trick here is to use [capacity-based compute model with Standard edition](https://docs.cloud.google.com/bigquery/docs/reservations-intro) instead of on-demand.  
This way instead of paying $6.25 over 35TB the query will cost $0.04 over 150 slotHr.

Test using sampled data before running full crawl to verify it with your billing account.

**Don’t run queries unless you’re confident about the expected cost!**

You’re correctly using all the clustering columns possible, but a partition filtering is required.  
Better test a single crawl before scanning more months in order to avoid high charge.

P.S. there is no retrospective HAR data in Cloud Storage anymore.
