Pythonic Crawling / Scraping Framework based on Non Blocking I/O operations.

Last update: Dec 05, 2022

Related tags

Web Crawling crawley

Overview

Pythonic Crawling / Scraping Framework Built on Eventlet

Features

High Speed WebCrawler built on Eventlet.
Supports relational databases engines like Postgre, Mysql, Oracle, Sqlite.
Supports NoSQL databased like Mongodb and Couchdb. New!
Export your data into Json, XML or CSV formats. New!
Command line tools.
Extract data using your favourite tool. XPath or Pyquery (A Jquery-like library for python).
Cookie Handlers.
Very easy to use (see the example).

Documentation

http://packages.python.org/crawley/

Project WebSite

http://project.crawley-cloud.com/

To install crawley run

~$ python setup.py install

or from pip

~$ pip install crawley

To start a new project run

~$ crawley startproject [project_name]
~$ cd [project_name]

Write your Models

""" models.py """

from crawley.persistance import Entity, UrlEntity, Field, Unicode

class Package(Entity):
    
    #add your table fields here
    updated = Field(Unicode(255))    
    package = Field(Unicode(255))
    description = Field(Unicode(255))

Write your Scrapers

""" crawlers.py """

from crawley.crawlers import BaseCrawler
from crawley.scrapers import BaseScraper
from crawley.extractors import XPathExtractor
from models import *

class pypiScraper(BaseScraper):
    
    #specify the urls that can be scraped by this class
    matching_urls = ["%"]
    
    def scrape(self, response):
                        
        #getting the current document's url.
        current_url = response.url        
        #getting the html table.
        table = response.html.xpath("/html/body/div[5]/div/div/div[3]/table")[0]
        
        #for rows 1 to n-1
        for tr in table[1:-1]:
                        
            #obtaining the searched html inside the rows
            td_updated = tr[0]
            td_package = tr[1]
            package_link = td_package[0]
            td_description = tr[2]
            
            #storing data in Packages table
            Package(updated=td_updated.text, package=package_link.text, description=td_description.text)


class pypiCrawler(BaseCrawler):
    
    #add your starting urls here
    start_urls = ["http://pypi.python.org/pypi"]
    
    #add your scraper classes here    
    scrapers = [pypiScraper]
    
    #specify you maximum crawling depth level    
    max_depth = 0
    
    #select your favourite HTML parsing tool
    extractor = XPathExtractor

Configure your settings

""" settings.py """

import os 
PATH = os.path.dirname(os.path.abspath(__file__))

#Don't change this if you don't have renamed the project
PROJECT_NAME = "pypi"
PROJECT_ROOT = os.path.join(PATH, PROJECT_NAME)

DATABASE_ENGINE = 'sqlite'     
DATABASE_NAME = 'pypi'  
DATABASE_USER = ''             
DATABASE_PASSWORD = ''         
DATABASE_HOST = ''             
DATABASE_PORT = ''     

SHOW_DEBUG_INFO = True

Finally, just run the crawler

~$ crawley run

Pythonic Crawling / Scraping Framework based on Non Blocking I/O operations.

Related tags

Overview

Pythonic Crawling / Scraping Framework Built on Eventlet

Features

Documentation

Project WebSite

To install crawley run

or from pip

To start a new project run

Write your Models

Write your Scrapers

Configure your settings

Finally, just run the crawler

Owner

Juan Manuel Garcia

EBay-email-tracker - Scapes an entire search page of a particular item on eBay and sends regular updates to an email address

Open Crawl Vietnamese Text

The core packages of security analyzer web crawler

此脚本为 python 脚本,实现原理为利用 selenium 定位相关元素,再配合点击事件完成浏览器的自动化.

Web Content Retrieval for Humans™

用python爬取江苏几大高校的就业网站，并提供3种方式通知给用户，分别是通过微信发送、命令行直接输出、windows气泡通知。

Simple Web scrapper Bot to scrap webpages using Requests, html5lib and Beautifulsoup.

A simple python script to fetch the latest covid info

A spider for Universal Online Judge(UOJ) system, converting problem pages to PDFs.

Basic-html-scraper - A complete how to of web scraping with Python for beginners

Telegram group scraper tool

Video Games Web Scraper is a project that crawls websites and APIs and extracts video game related data from their pages.

Meme-videos - Scrapes memes and turn them into a video compilations

Examine.com supplement research scraper!

Ebay Webscraper for Getting Average Product Price

Scrape plants scientific name information from Agroforestry Species Switchboard 2.0.

Complete pipeline for crawling online newspaper article.

Python script for crawling ResearchGate.net papers✨⭐️📎

🥫 The simple, fast, and modern web scraping library

Use Flask API to wrap Facebook data. Grab the wapper of Facebook public pages without an API key.