andyzsf /
pyspider
A Powerful Spider(Web Crawler) System in Python. http://docs.pyspider.org/
37/100 healthLoading repository data…
binux / repository
A Powerful Spider(Web Crawler) System in Python.
A transparent discovery signal based on current public GitHub metadata.
This score does not audit code, security, maintainers, documentation quality, or suitability. Verify the repository and its current documentation before adoption.
A Powerful Spider(Web Crawler) System in Python.
Tutorial: http://docs.pyspider.org/en/latest/tutorial/
Documentation: http://docs.pyspider.org/
Release notes: https://github.com/binux/pyspider/releases
from pyspider.libs.base_handler import *
class Handler(BaseHandler):
crawl_config = {
}
@every(minutes=24 * 60)
def on_start(self):
self.crawl('http://scrapy.org/', callback=self.index_page)
@config(age=10 * 24 * 60 * 60)
def index_page(self, response):
for each in response.doc('a[href^="http"]').items():
self.crawl(each.attr.href, callback=self.detail_page)
def detail_page(self, response):
return {
"url": response.url,
"title": response.doc('title').text(),
}
pip install pyspiderpyspider, visit http://localhost:5000/WARNING: WebUI is open to the public by default, it can be used to execute any command which may harm your system. Please use it in an internal network or enable need-auth for webui.
Quickstart: http://docs.pyspider.org/en/latest/Quickstart/
Licensed under the Apache License, Version 2.0
Selected from shared topics, language and repository description—not editorial ratings.
andyzsf /
A Powerful Spider(Web Crawler) System in Python. http://docs.pyspider.org/
37/100 healthTMLoew /
A Powerful Spider(Web Crawler) System in Python.
47/100 healthMichael0711 /
A Powerful Spider(Web Crawler) System in Python.
40/100 healths4dman /
A Powerful Spider(web crawler) System built in Python.
27/100 healthHenryKamg /
A Powerful Spider(Web Crawler) System in Python. http://docs.pyspider.org/
37/100 health