AI & ML interests

A one-year long research workshop on large language models: the Summer of Language Models 21 🌸

Recent Activity

BramVanroy 
posted an update about 1 month ago
view post
Post
2734
**I benchmarked HF buckets against https access for Common Crawl.**

Took me a while to get round to do this but I benchmarked access to Common Crawl via https vs hf buckets. Both experiments were run at night in Europe. I do not think other hardware problems were impacting the speeds since CPU processing time of the non-download pipeline components were highly similar (within 2% identical) and below only the WarcReader speeds of datatrove are used.

Experiment: selected 5 disjoint samples of 64 files each (randomly from the latest crawl; 20,499 docs/file). Those five batches were then processed by 32 single-core tasks with 4GB/core (five batches to calculate CIs). Paired experiment between using https and hf bucket.

- https: 40.0 [39.3-40.6] (seconds per WARC file)
- hf bucket: 172.0 [122.7-221.2]

That is a difference of about 4x in streaming speed. You'll see that https is also more stable (smaller CI).

I also ran raw throughput tests to the endpoints to measure rate limiting (64MiB transfer at 8/32/128/256 concurrent readers) and rate limiting seems not an issue for either: at any of those parallel reader numbers, their respective speeds stay about the same.

Note that, given CC scale, this is still a small test. Rate limiting may become more obvious when processing a full crawl. I do not know whether the https endpoint vs HF bucket will shut you out earlier with which limits.
christopher 
in bigscience/bloom about 2 months ago

Fixed CO2 metadata in README.md

1
#293 opened about 2 months ago by
Croc-Prog-HF
julien-c 
posted an update about 2 months ago
view post
Post
5350
who's working on an NVFP4 version of Kimi-K3?
  • 4 replies
·