Anywhere: A Web Crawler Automation Management Interface
Abstract
Web crawling projects or design is significant in the current information age. Using the web spider or crawler can automatically search and collect a huge amount of internet information. As one of the most popular web crawler frameworks, Scrapy is robust in abundant functions but weak in easy operation. In this paper, we provide a framework Anywhere, for optimising the usage feeling and improving the use efficiency of the web crawling management of Scrapy. We analyse the whole workflow of a web crawling project of Scrapy and design two main functions in Anywhere, one is quickly generating a Scrapy project with the preset temperatures, the other is repeatable configuration function for the Scrapy project setting. Beside, with Anywhere, users can easily directly manage multiple Scrapy projects with a file folders architecture. Compared with normal Scrapy project interactive coding development, we test Anywhere with enough experiments that show Anywhere can improve the development efficiency of Scrapy projects to about 200\%. For the multiple project management in code interaction level, the developing efficiency is improved to about 300\%. We simplify the procedure to quickly generate a simple spider project with Scrapy. Anywhere can assist the development of Scrapy is useful for the design of large batch concurrent projects at coding level.
- Publication:
-
arXiv e-prints
- Pub Date:
- May 2024
- DOI:
- 10.48550/arXiv.2407.00025
- arXiv:
- arXiv:2407.00025
- Bibcode:
- 2024arXiv240700025L
- Keywords:
-
- Computer Science - Distributed;
- Parallel;
- and Cluster Computing
- E-Print:
- 8 pages