Collecting Raw Data
Gathering play-by-play, player stats, and betting lines from varied public and commercial sources.
The analytical framework described here serves to systematically examine basketball players, teams, and betting markets using empirical data. It encompasses data from box scores, which list basic statistics like points, rebounds, and assists per game, as well as play-by-play logs that capture every event in sequence. Historical archives provide long-term records, enabling contextual comparisons across eras. The general principles involve quantifying individual and team performance through metrics such as efficiency ratings, pace-adjusted statistics, and on/off court differentials. For betting lines, the framework assesses probabilities by analyzing historical spreads and outcomes. The scope is bounded to descriptive and inferential analysis, not predictive modeling. No advice or forecasts are implied.
Gathering play-by-play, player stats, and betting lines from varied public and commercial sources.
Standardizing formats, removing errors, and merging datasets into a consistent structure for analysis.
Applying statistical models to assess performance, trends, matchups, and lineup dynamics across contexts.
Comparing market expectations with model outputs to identify conditional deviations, dependent on external factors.
Hoops Archive serves as a repository for basketball data and the methods used to analyze it. The site collects statistics from games, player histories, and betting lines, then organizes this information into contextual archives. We focus on documenting the analytical approaches applied to these datasets—such as regression models, pace-adjusted metrics, or historical comparisons—so readers can see how evaluations are constructed. The material is informational, not predictive, and acknowledges that different datasets and models may yield varying interpretations. Our goal is to support individual research and discussion among fans, analysts, and bettors.
When evaluating basketball performance and betting lines, the reliability of conclusions depends heavily on data quality, source transparency, and methodological awareness. Different data providers may record events inconsistently, and various analytical approaches—such as per-possession metrics or raw totals—can lead to varying interpretations. Discrepancies in sample sizes, updating frequency, and contextual factors further affect how numbers are read. Being explicit about these elements helps readers understand the limitations and proper application of any analysis.