Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
admp
on Oct 26, 2012
|
parent
|
context
|
favorite
| on:
80 terabytes of archived web crawl data available ...
CommonCrawl also has a fairly large ("The crawl currently covers 5 billion pages") dataset of this sort, which unlike the one from archive.org is already available to everyone on S3 under the requester-pays model.
http://commoncrawl.org/data/accessing-the-data/
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search:
http://commoncrawl.org/data/accessing-the-data/