FIFA-World-Cup-2026-Dataset GitHub Details, Stars and Alternatives | OpenRepoFinder
mominullptr / repository
FIFA-World-Cup-2026-Dataset
The most complete and actively updated FIFA World Cup 2026 dataset on Kaggle. A clean, authentic, relational dataset covering the entire tournament (June 11 – July 19, 2026) with the first-ever 48-team format. Includes real match results updated daily, 1,248 players across all 48 squads, expected goals (xG), minute-by-minute match events.
A transparent discovery signal based on current public GitHub metadata.
Recent activity35% weight
100
Community adoption25% weight
30
Maintenance state20% weight
100
License clarity10% weight
0
Project information10% weight
100
This score does not audit code, security, maintainers, documentation quality, or suitability. Verify the repository and its current documentation before adoption.
README preview
FIFA World Cup 2026 Dataset- Live & Updated Stats — Matches, Squads, Players, Stats & Live Results
The most complete and actively updated FIFA World Cup 2026 dataset on Kaggle. A clean, authentic, relational dataset covering the entire tournament (June 11 – July 19, 2026) with the first-ever 48-team format. Includes real match results updated daily, 1,248 players across all 48 squads, expected goals (xG), minute-by-minute match events, and per-team match statistics sourced from FIFA.com and verified providers.
Why this dataset? Unlike other World Cup datasets, this one is (1) updated daily with real match results as games are played, (2) fully relational with normalized foreign keys for SQL/database modeling, (3) contains zero synthetic data — every stat is sourced and traceable, and (4) includes advanced metrics like xG, match team stats, and VAR events.
🏆 Tournament Status
Tournament: FIFA World Cup 2026
Format: 48 Teams
Current Stage: Quarter-finals
Matches Completed: 96
Updated After: Every Completed Match
Validation: Relational integrity verified before every release
👥 Community Recognition
The FIFA World Cup 2026 Dataset has started gaining traction within the data science and developer community.
Mentioned by Software with Nick in an Instagram reel highlighting useful datasets for AI, machine learning, and data science projects.
Featured in the Analytics Jobs in Sports newsletter on Substack, highlighting the dataset for sports analytics.
Featured in a LinkedIn post that reached 22,000+ professionals, generating significant engagement from the data community.
Shared and discussed by developers building World Cup prediction models, dashboards, SQL projects, and sports analytics applications.
ALGORITHMICALLY RELATED
Similar Open-Source Projects
Selected from shared topics, language and repository description—not editorial ratings.
Sri Venkateshwara University (SVU) strives to create professionals who are not only adept in academics but also in application for the benefit of humanity. We foster a culture of learning by doing. We believe in nurturing students who are at the forefront of innovation by offering an environment of research & development to make us Best University in Uttar Pradesh (UP). SVU believes in experiential learning. To facilitate this, we have an ultra-modern infrastructure that motivates students to experiment & excel in their area of interest. The Best University of Moradabad has laboratories & workshops that signify our commitment to core research, thus enabling innovation. SVU is the only institution to have set up labs in collaboration with the industry. This way we can train our students on the latest skills & make them employable. Students sharpen their practical skills under the watch full eyes of trainers & become competent professionals. For the overall development of the students, we organize cultural programs. Students take part in these programs & exhibit their talent to become confident professionals. The annual fest attracts students from all over the country & showcase their talent to make us the Top University in India. We equipped the computing labs with the latest software & hardware to augment the technical skills of the students. SVU’s library is an epitome of knowledge. It has over 3000 books & journals that ensure the students are never short on intellectual input. The team of industry trainers educate them on the key skills so crucial for employment & make us the Best University in Gajraula. The specially created engineering labs assist engineers to refine their technical acumen so much needed for the country. The Chairman Dr. Sudhir Giri believes in removing all the economic & social barriers that can hinder education. Hence, SVU provides many scholarships & grants to meritorious students. Up till now, the college has enabled over 500000 students to attain their academic desires to make us the Best Private University in Uttar Pradesh (UP). The group is running a dozen educational institutions that include medical colleges in India & abroad. Our commitment towards education & healthcare has enabled Dr Sudhir Giri to win the International Glory Man of the year Award 2021. The Best Private University in Moradabad is on the Delhi Moradabad highway, well connected with rail & road. The green surroundings provide peace of mind that enables research based learning. The carefully recruited faculty is the pride of the university. They have years of industrial & academic experience so vital for the students. They transfer key skills & make us the Best Private University in Gajraula. The faculty encourages students to undertake research & sharpen their skills that will enable them to get jobs. Majority of the faculty members are doctorates who educate the students to become competent professionals. The faculty takes part in FDP in order to develop a culture of research. The specialty of SVU is the internship. We have partnered with leading industries for providing internship to the students. We believe that education without applicability is incomplete. Students gain hands on exposure through internship & become job ready. We place most of the students during internship to make us the Top University in India. SVU, the Best University in Uttar Pradesh (UP), adopts a futuristic teaching pedagogy. We strive for experiential learning of our students through role plays, projects & presentation. The students take part in the learning activity & imbibe concepts that enable their placements. The AC seminar & conference halls allow knowledge dispersion for the development of the students. The University is running over 150 undergraduate (UG), postgraduate (PG) courses, (Ph.D.), diploma and certificate courses in various fields of Applied Sciences, Medical Science, Humanities & Social Sciences. We also run courses in Languages, Design, Agriculture, Engineering & Technology, Nursing, Pharmacy, Paramedical, Commerce & Management, Law, Library & information Sciences, Mass Comm. & Journalism to enhance the employability of the youth. SVU has a culture of project based learning. Students do projects in each semester under the guidance of faculty. They complete these projects in earmarked industries to garner hands-on skills. Through these projects, we train students on the hot skills so crucial for employment to make us the Best University in Moradabad. SVU’s Research & Development (R&D) wing encourages students to work on research areas important for the country. We have partnered with leading research institutions to undertake research. The breath-taking infrastructure of the best university in Gajraula motivates researchers to achieve their goals for research. Owing to our dedication, SVU has received grants from GOI for research on areas of national importance. The faculty members provide guidance to the scholars until they achieve their aim. We have set up the incubation center to provide fillip to new ideas that foster entrepreneurship. We want to be an institution that supports the ‘Make in India’ vision of the government. The center supports new ideas that enable the young entrepreneurs to create startups & become successful. Under the strong leadership of Dr. Sudhir Giri, till date we have successfully incubated 150 start-ups. This speaks of our exemplary education & make us the Best Private University in Uttar Pradesh (UP). These startups are not only creating wealth but also providing employment to the needy. The industrialists have lamented that the epicenter for entrepreneurship will be the educational institutions. We need to provide them with the support & infrastructure for this. The annual hackathon attracts individuals who showcase their business acumen to make us the Best Private University in Moradabad. SVU has a dedicated International Research & collaboration Cell (IRCC) that collaborates with universities abroad. Faculty & students who want to pursue studies abroad the IRCC starts admission formalities for them. We have partnered with reputed institutions for providing excellent research collaborations. Those who wish to do P. HD abroad the IRCC help them gain admission & make us the Top University in India. A lot of our faculty members are pursuing their research internationally & contributing to the welfare of humanity. SVU strives to make our students feel comfortable at the campus. Separate hostel for boys & girls with 24 hour security is available at SVU. The cafeteria serves nutritious food to the students. Gym, recreation hall & the sports ground help to relax our students & make us the Best University in Uttar Pradesh (UP). The campus has an in house ATM & convenience store for the benefit of the students. SVU enables placement through exemplary training. We train on communication & interpersonal skills in order to refine the personality of the students. We make them practice mock interviews & group discussion that help to clear placement tests. Ninety percent of the students get placed before their last semester to make us the best university in Moradabad. We have hired industrial trainers in order to provide training on block chain, machine learning, artificial intelligence (AI), and python & data science. These trainers have years of experience that enables them in training the students. The students gain key insights on these technologies & sharpen their acumen to make us the Best University in Gajraula.
Continuously updated after every completed FIFA World Cup 2026 match.
Real-World Group Configurations: Reflects the actual 12 groups (Groups A to L) with zero qualifiers placeholders.
Geographical & Altitude Details: Includes coordinates (Latitude/Longitude) and exact elevations in meters for all 16 host stadiums in the USA, Canada, and Mexico.
Granular Player Data: Features 1,248 players (26-man squads for all 48 teams) with market values in Euros, national team caps, positions, and club teams.
Live Expected Goals (xG): Includes expected goals metrics for all completed matches.
Match Events: Logs every goal, assist, card, and VAR review chronologically by minute.
Match Team Stats: Per-team per-match statistics (possession, shots, corners, etc.) sourced from verified providers.
Interactive Updater: Built-in interactive console application to easily record daily match outcomes and populate events without breaking database normalization.
ML-Ready Feature Set: A curated dataset with 65 pre-calculated features (ELO ratings, FIFA rankings, squad market values, rolling team form, and fatigue) ready to train machine learning classifiers.
🤖 ML-Ready Match Prediction Dataset
Skip the tedious feature engineering. This repository includes match_prediction_features.csv, a dataset engineered specifically for sports analytics and machine learning:
import pandas as pd
from xgboost import XGBClassifier
# Load dataset
df = pd.read_csv("match_prediction_features.csv")
# Train on completed matches (1 to 100), predict upcoming (101+)
train_df = df[df["match_id"] <= 100]
predict_df = df[df["match_id"] > 100]
# Features: Elo differences, Squad value ratio, Rolling form xG
features = ["home_elo", "away_elo", "home_fifa_rank", "away_fifa_rank",
"home_squad_total_value_eur", "home_prev_avg_xg_scored", "away_prev_avg_xg_scored"]
# Fit XGBoost Classifier to predict match result ('H', 'D', 'A')
clf = XGBClassifier()
clf.fit(train_df[features], train_df["match_result"])
# Predict upcoming knockout stages (Semi-Finals & Final)
predictions = clf.predict(predict_df[features])
print("Forecasted outcomes:", predictions)
Database Schema
erDiagram
TEAMS {
int team_id PK
string team_name
string fifa_code
string group_letter
string confederation
int fifa_ranking_pre_tournament
int elo_rating
string manager_name
}
VENUES {
int venue_id PK
string stadium_name
string city
string country
int capacity
float latitude
float longitude
int elevation_meters
}
TOURNAMENT_STAGES {
int stage_id PK
string stage_name
bool is_knockout
}
REFEREES {
int referee_id PK
string name
string country
float avg_cards_per_game
}
MATCHES {
int match_id PK
string date
string kickoff_time_utc
int stage_id FK
int venue_id FK
int home_team_id FK
int away_team_id FK
int home_score
int away_score
string status
float home_xg
float away_xg
int referee_id FK
}
SQUADS_AND_PLAYERS {
int player_id PK
int team_id FK
string player_name
string position
string club_team
int market_value_eur
int caps
string date_of_birth
int height_cm
int goals
}
MATCH_EVENTS {
int event_id PK
int match_id FK
int minute
string event_type
int team_id FK
int player_id FK
}
MATCH_LINEUPS {
int lineup_id PK
int match_id FK
int player_id FK
int team_id FK
int is_starting_xi
string tactical_position
int minutes_played
}
PLAYER_STATS {
int player_id PK, FK
string player_name
int team_id FK
int matches_played
int matches_started
int minutes_played
int goals
int assists
int yellow_cards
int red_cards
int penalty_scored
int own_goals
int saves
int goals_conceded
int clean_sheets
string data_source
string last_verified
}
TEAMS ||--o{ MATCHES : "hosts/visitors"
TEAMS ||--o{ SQUADS_AND_PLAYERS : "roster"
VENUES ||--o{ MATCHES : "hosts"
TOURNAMENT_STAGES ||--o{ MATCHES : "stage"
REFEREES ||--o{ MATCHES : "officiates"
MATCHES ||--o{ MATCH_EVENTS : "contains"
SQUADS_AND_PLAYERS ||--o{ MATCH_EVENTS : "triggers"
MATCHES ||--o{ MATCH_LINEUPS : "lineups"
SQUADS_AND_PLAYERS ||--o{ MATCH_LINEUPS : "plays_in"
TEAMS ||--o{ MATCH_LINEUPS : "lineups"
SQUADS_AND_PLAYERS ||--|| PLAYER_STATS : "has"
TEAMS ||--o{ PLAYER_STATS : "has"
MATCH_TEAM_STATS {
int match_id FK
int team_id FK
int possession_pct
int total_shots
int shots_on_target
int corners
int fouls
int offsides
int saves
string data_source
string last_updated
}
MATCHES ||--o{ MATCH_TEAM_STATS : "stats"
TEAMS ||--o{ MATCH_TEAM_STATS : "stats"
CSV Files Description
teams.csv: Information on all 48 participating countries.
venues.csv: Geolocation, capacities, and elevation details of all 16 stadiums.
tournament_stages.csv: Lookup table for stages (Group Stage, Round of 32, etc.).
referees.csv: International referees with their historical card-per-game stats.
matches.csv: Match outcomes, dates, times, scores, xG metrics, and statuses using relational IDs (stage_id, venue_id, etc.) for clean database modeling.
matches_detailed.csv: A denormalized, user-friendly version of matches.csv that displays human-readable names (e.g. home_team_name, stadium_name, city, referee_name) instead of IDs. Ideal for quick analysis without SQL joins!
squads_and_players.csv: Detailed player registries (1,248 rows) containing verified player names (preserved with native accents), positions, clean club teams, market values, international caps, dates of birth (in YYYY-MM-DD format), heights in centimeters, and international goals.
match_events.csv: Time-series game events (goals, assists, cards, VAR reviews) mapped to matches and players.
match_team_stats.csv: Per-team per-match statistics (possession %, shots, shots on target, corners, fouls, offsides, saves) with data_source and last_updated columns for full traceability. Only populated with verified data from authentic sources (FIFA, Sofascore, FBref, etc.).
match_lineups.csv: Tactical lineups for all completed matches: starting XI (11 players per team) and substitutes with actual minutes played.
player_stats.csv: Cumulative tournament statistics for each player (1,248 rows), updated continuously as matches conclude. Outfield players have goalkeeper-specific fields set to NULL, and unverified advanced metrics (such as shots or key passes) are kept NULL to preserve dataset authenticity.
match_prediction_features.csv: A machine-learning-ready features dataset containing 65 pre-calculated predictive features compiled for every match of the World Cup (completed and upcoming knockout stages). Features include team ratings (ELO, FIFA ranking, squad value), rolling team form (possession, shots, goals, xG, corners, saves), player fatigue (rest days), environmental factors (stadium elevation), and final target labels. Load directly into pandas or scikit-learn for rapid model training.
📋 Data Integrity Policy
Hard rule: No synthetic or generated match stats, assists, or events are added to public tables.
All match results, events, and statistics are sourced from verified, authentic providers.
New fields (e.g., assists, team stats) are only populated when confirmed from real sources.
The data_source column in match_team_stats.csv provides full traceability.
If verified data is unavailable for a match, that match is simply omitted from optional tables rather than filled with generated values.
🛠️ Installation & Usage
To generate the initial CSV dataset or regenerate it back to its default start state:
python generate_dataset.py
Daily Match Updates
As matches conclude every day, you can update the datasets interactively. The script will guide you step-by-step to record final scores, xG, and select players for goals, cards, and VAR reviews:
python update_dataset.py
SQLite Database Generation
For researchers and SQL query design, you can package all normalized CSV files into a single, query-optimized SQLite relational database file:
python generate_sqlite.py
🏷️ Citation
If you use this dataset in your research, publications, or projects, please cite it using the academic metadata provided in CITATION.cff or using the format below:
@dataset{fifa_world_cup_2026,
author = {MD Mominul Islam},
title = {FIFA World Cup 2026 Dataset- Live & Updated Stats}
In this course, you’ll use Python to refactor an existing application that was originally built using Node.js. The app, called Just Tech News, is a website where users can post, upvote, and comment on links to news articles. Instead of using Node.js as the back-end language and Handlebars.js as the templating engine, you’ll use Python to create the app and the Python Flask framework to create the application’s views. Python is one of the most widely used high-level programming languages, with the TIOBE index naming it the second most popular language out of 100 in 2020. It's an interpreted, high-level, open-source, general-purpose programming language that supports procedural, object-oriented, and some functional programming constructs. Created by Guido van Rossum between 1985 and 1990, Python is currently used in such fields as web development, data science, and DevOps. Python has seen an increase in popularity with the rise of big data—in which extremely large datasets are analyzed computationally to reveal patterns, trends, and associations, especially relating to human behavior and interactions. Python can help us mine big data much faster than any other programming language because it can handle any data type. Python can also be easily integrated into web applications that need to implement machine learning. Machine learning is a quickly growing branch of artificial intelligence that's based on the idea that systems can learn from data, identify patterns, and make decisions with minimal human intervention. Just like most other popular programming languages, Python has several frameworks that can make it easier to use. The two most popular Python frameworks are Flask and Django. You’ll use Flask in this course, but you can adapt the same concepts to a Django application. To complete the refactoring, you’ll connect the Just Tech News application to a relational database using SQLAlchemy, provide user authentication using Flask’s built-in session functionality, and deploy your app to the cloud using Heroku.