Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Curious how useful very large video datasets are compared with smaller, better curated ones.
At some point, adding more videos probably helps less than improving the quality of the data.
by techsage - - can someone with expertise give us an overview of the architecture involved doing this
- let us say you ran yt-dlp inside python aiohttp
- surely your ll run a limit soon as your ip address will be flagged
- what solutions do we have to auto rotate proxies in python
- are there better, faster and more reliable ways to go about downloading a 100 million videos without getting your ip address blocked?
by vivzkestrel - Damn, that's a lot of videos for them to contact the creators and ask for permission to use their content as machine learning training content. Unless of course, they didn't, and just went ahead with it anyway...by voidUpdate
- "Overall, we attempted to download 130M videos and achieved a link success rate of approximately 60%, resulting in 80M successfully retrieved videos with a total duration of 10M hours."
I am astonished that the success rate is so high. How Youtube didn't block them, I don't know. But I think that this URL list won't age well because youtube will very quickly block any researcher trying to download these videos themselves.
by topwalktown