GoogleSpider

Crawl the information of a given keyword on Google search engine

Config

DataBase

Currently, data is stored in mongodb, and the database configuration is in line 15-19 of the setting. py file, which can be modified by yourself.

# MONGODB
MONGO_IP = "localhost"
MONGO_PORT = 27017
MONGO_DB = "Google_spider"
MONGO_USER_NAME = ""
MONGO_USER_PASS = ""

Log

LOG_NAME = os.path.basename(os.getcwd())
LOG_PATH = "log/%s.log" % LOG_NAME  # log path
LOG_LEVEL = "DEBUG"
LOG_COLOR = True  
LOG_IS_WRITE_TO_CONSOLE = True 
LOG_IS_WRITE_TO_FILE = True  
LOG_MODE = "w" 
LOG_MAX_BYTES = 10 * 1024 * 1024  # Maximum bytes
LOG_BACKUP_COUNT = 20  # Number of log files reserved
LOG_ENCODING = "utf8"  # code
OTHERS_LOG_LEVAL = "ERROR"  # leval

Spider

Download interval
- ```
SPIDER_SLEEP_TIME = [0, 1]
```
Maximum number of requests (100 by default)
- ```
SPIDER_MAX_RETRY_TIMES = 100
```
  Note
  
  If an illegal interface is encountered during crawling, an exception of 'user agent -- illegal interface' will be thrown, and then the crawler task will retry until the data is successfully crawled or more than 100 times

data structure

key	value type	example
title	str	“Donald Trump - Wikipedia”
keyword	str	“Trump"
url	str	"https://en.wikipedia.org/wiki/Donald_Trump"
text	str	Donald Trump - Wikipedia 1 hour ago · Donald John Trump (born June 14, 1946) is an American politician, media personality, and businessman who served as the 45th president of the United States ... Vice President: Mike Pence In office January 20, 2017 – January 20, 2021: In office; January 20, 2017 – January 20, 2021 Occupation: Politician; businessman; television presenter Parents: Fred Trump; Mary Anne MacLeod"

Quick start

Crawl the 3 page data with the keyword 'Trump'

from spiders.google_curl import GoogleCurl

spider = GoogleCurl('Trump', 3)
spider.start()

The first parameter is the search keyword, and the second parameter is the number of pages crawled

Crawl the information of a given keyword on Google search engine

Related tags

Overview

GoogleSpider

Config

DataBase

Log

Spider

data structure

Quick start

Owner

The first public repository that provides free BUBT website scraping API script on Github.

This Scrapy project uses Redis and Kafka to create a distributed on demand scraping cluster

基于Github Action的定时HITsz疫情上报脚本，开箱即用

🐞 Douban Movie / Douban Book Scarpy

A Smart, Automatic, Fast and Lightweight Web Scraper for Python

Twitter Claimer / Swapper / Turbo - Proxyless - Multithreading

A web scraping pipeline project that retrieves TV and movie data from two sources, then transforms and stores data in a MySQL database.

Docker containerized Python Flask API that uses selenium to scrape and interact with websites

Python script that reads Aliexpress offers urls from a Excel filename (.csv) and post then in a Telegram channel using a bot

爬取各大SRC当日公告 | 通过微信通知的小工具 | 赏金工具

Scrapy-soccer-games - Scraping information about soccer games from a few websites

Nekopoi scraper using python3

爬虫案例合集。包括但不限于《淘宝、京东、天猫、豆瓣、抖音、快手、微博、微信、阿里、头条、pdd、优酷、爱奇艺、携程、12306、58、搜狐、百度指数、维普万方、Zlibraty、Oalib、小说、招标网、采购网、小红书》

Scrape puzzle scrambles from csTimer.net

Github scraper app is used to scrape data for a specific user profile created using streamlit and BeautifulSoup python packages

Scrape plants scientific name information from Agroforestry Species Switchboard 2.0.

Scraping weather data using Python to receive umbrella reminders

Script used to download data for stocks.

An utility library to scrape data from TikTok, Instagram, Twitch, Youtube, Twitter or Reddit in one line!

Minecraft Item Scraper