Vacancy Description
About the Company
We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models. We power and support Grass, a bandwidth‑sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale.
Responsibilities
- Build and maintain large-scale web crawlers across diverse domains
- Design high-throughput, fault-tolerant systems for data collection
- Handle anti-bot systems, rate limits, and dynamic/JS-heavy sites
- Develop pipelines for cleaning, deduplication, filtering, and normalization
- Construct and maintain datasets for research and model training
- Monitor crawl performance, coverage, and data quality
- Collaborate with research teams to align data collection with modeling needs
- Optimize infrastructure for cost, latency, and reliability
Ready to Apply?
अभी आवेदन करें
Submit your application for Research Crawling Engineer at Remotedxb
Apply for this Position