REPOSITORY OVERVIEWLive repository statistics
★ 0Stars
⑂ 0Forks
◯ 0Open issues
◉ 0Watchers
27/100
OPENREPOHUB HEALTH SIGNALLimited signals
A transparent discovery signal based on current public GitHub metadata.
Recent activity35% weight
10 Community adoption25% weight
0 Maintenance state20% weight
100 License clarity10% weight
0 Project information10% weight
35 This score does not audit code, security, maintainers, documentation quality, or suitability. Verify the repository and its current documentation before adoption.
README preview
This project serves as a comprehensive guide to building an end-to-end data engineering pipeline. It covers each stage from data ingestion to processing and finally to storage, utilizing a robust tech stack that includes Apache Airflow, Python, Apache Kafka, Apache Zookeeper, Apache Spark, and Cassandra. Everything is containerized using Docker for ease of deployment and scalability.
The workflow of the project

The project is designed with the following components:
Project Overview
This project leverages a variety of technologies to establish a robust data pipeline. Below is an outline of the key components and their roles within the system.
Data Source
- Random User Generator: We utilize the randomuser.me API to generate random user data, which serves as the initial input for our pipeline.
Data Orchestration
- Apache Airflow: Orchestrates the pipeline, managing the workflow from data ingestion to storage. Airflow is responsible for fetching data from the Random User Generator API and storing it in a PostgreSQL database.
Data Streaming
- Apache Kafka and Zookeeper: Facilitate the streaming of data from the PostgreSQL database to our processing engine. Kafka, in conjunction with Zookeeper, ensures reliable and scalable streaming capabilities.
Monitoring and Schema Management
- Control Center and Schema Registry: Essential for the monitoring and management of Kafka streams. The Control Center provides a comprehensive overview of the system's health and performance, while the Schema Registry maintains the integrity of data schemas within our streams.
Data Processing
- Apache Spark: Employs both master and worker nodes to process the data efficiently. Apache Spark's robust computing capabilities enable complex data transformations and analyses.
Data Storage
- Cassandra: The final destination for our processed data. Cassandra's distributed architecture ensures scalability and reliability for storing large volumes of data.
ALGORITHMICALLY RELATEDSimilar Open-Source Projects
Selected from shared topics, language and repository description—not editorial ratings.
The Voting System web application using Django is a project that serves as the automated voting system of an organization or school. This system works like the common manual system of election voting system whereas this system must be populated by the list of the positions, candidates, and voters. This system can help a certain organization or school to minimize the voting time duration because aside they can provide the voters an online platform to vote, the system will automatically count the votes for each candidate. The system has 2 sides of the user interface which are the administrator and voters side. The admin user is in charge to populate and manage the data of the system and the voter side which is where the voters will choose their candidate and submit their votes.
81/100 healthRecently updatedActive repository
PythonMIT#django#django-project#e-voting#otp
⑂ 86 forks◯ 3 issuesUpdated 4 days ago
This project serve HTML files (and a few more) saved in your computer with a UI suitable for Kindle web browser. On top of that, it include a Read Mode (thanks to ReadabiliPy) to display the text in a comfortable size without have to use the 'Article Mode' in Kindle web browser.
62/100 healthActive repository
PythonNo license#flask