Tracking comparative Enterprise Linux usage primarily using EPEL usage metrics https://rocky-stats.tiuxo.com
  • Python 46.1%
  • Jupyter Notebook 30%
  • HTML 23.9%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
Brian Clemens c51eb6e825
All checks were successful
Update statistics / notebooks (push) Successful in 1m18s
Serialize stats runs and avoid needless cache transfers
Two concurrent runs both pulling the ~600MB cache object wedged the S3
gateway, so add a concurrency group. Also compare ETags using only the
small sidecar object and pull the cached CSV only when it can actually
be reused, instead of always transferring it and sometimes discarding
it in favour of an upstream download. Cap AWS retries so a stalled
endpoint fails fast rather than hanging for minutes.
2026-08-12 16:18:57 +09:00
.forgejo/workflows Serialize stats runs and avoid needless cache transfers 2026-08-12 16:18:57 +09:00
public Move web pages into public/ and update source link to Forgejo 2026-07-19 22:01:04 +09:00
.gitignore gitignore 2025-08-14 20:03:26 +09:00
config.py Count each host once via its primary EPEL repo 2026-08-12 15:57:38 +09:00
data_utils.py Count each host once via its primary EPEL repo 2026-08-12 15:57:38 +09:00
epel.ipynb Fix chart resample title 2025-12-11 19:00:24 +09:00
LICENSE Initial commit 2022-05-15 12:49:32 +09:00
plot_utils.py refactor 2025-08-14 19:54:22 +09:00
README.md Count each host once via its primary EPEL repo 2026-08-12 15:57:38 +09:00
requirements.txt Remove unused download machinery after CI-side data fetch 2026-07-21 18:57:42 +09:00

Rocky Stats - Enterprise Linux Distribution Analysis

A data visualization project analyzing Enterprise Linux distribution usage patterns using EPEL telemetry and DockerHub pull statistics.

Overview

This project generates comprehensive charts comparing Rocky Linux, RHEL, CentOS variants, AlmaLinux, and Oracle Linux usage trends over time. The analysis helps understand the Enterprise Linux ecosystem evolution and adoption patterns.

Features

  • Automated Data Collection: Downloads latest EPEL and DockerHub data with smart caching
  • Distribution Analysis: Tracks market share and usage trends across major EL distributions
  • Architecture Breakdown: Analyzes usage across x86_64, aarch64, ppc64le, and s390x architectures
  • System Age Analysis: Differentiates between ephemeral and long-term deployments
  • Version Tracking: Monitors adoption of different EL major/minor versions
  • Export Formats: Generates both PNG and SVG outputs for presentations and web use

Project Structure

├── config.py              # Configuration settings and constants
├── data_utils.py           # Data loading and preprocessing functions  
├── plot_utils.py           # Plotting and visualization utilities
├── epel_refactored.ipynb   # Main analysis notebook (refactored)
├── epel.ipynb              # Original notebook (legacy)
├── requirements.txt        # Python dependencies
├── data/                   # Downloaded data files (cached)
├── out/                    # Generated chart outputs
└── LICENSE                 # MIT License

Setup

  1. Install Dependencies:

    pip install -r requirements.txt
    
  2. Provide Input Data: The notebook reads data/epel.csv directly and does not download it. Download the EPEL countme totals from the source URL (see config.EPEL_DATA_URL) into the data/ directory before running:

    mkdir -p data
    curl -o data/epel.csv \
      https://data-analysis.fedoraproject.org/csv-reports/countme/totals.csv
    
  3. Run Analysis:

    jupyter notebook epel.ipynb
    
  4. Output Location: Generated charts are saved to the out/ directory in both PNG and SVG formats.

Data Sources

In CI the EPEL data is fetched by the workflow and cached in object storage, re-downloading from the source only when it changes (via HTTP ETag).

How systems are counted

The countme data records one hit per repository per host per week, not one per host. A host with several EPEL repos enabled therefore appears several times. To approximate a per-host count, only the primary repos listed in config.EPEL_REPOS are counted, and a repo's major version must match the host's. Supplemental repos enabled alongside a primary one (EPEL Next, modular, testing) are excluded.

This matters most for CentOS Stream: EPEL Next exists specifically for Stream, around three quarters of Stream hosts enable it, and Stream accounts for ~87% of all EPEL Next traffic. Counting it would roughly double Stream's totals.

Charts therefore show hosts counted once via their primary repo. True unique instance counts are not recoverable from this data.

Only released (GA) versions are counted; see the notes on the version lists in config.py.

Generated Charts

Distribution Analysis

  • Enterprise Linux instances by distribution (share and total)
  • Long-term trends with polynomial trendlines
  • Ephemeral vs. persistent deployments

Version-Specific Analysis

  • EL8, EL9, and EL10 adoption patterns
  • Rocky Linux version distribution over time

Architecture Analysis

  • AltArch (non-x86_64) usage patterns
  • Rocky Linux deployment by architecture

System Age Analysis

  • Distribution breakdown by system age categories
  • Rocky Linux deployment maturity patterns

Configuration

Key settings in config.py:

  • EMPHASIZE: Which distribution to highlight in charts (default: 'Rocky Linux')
  • COLORS: Brand colors for each distribution
  • STARTDATE: Analysis time window (default: 1 year)

Development

Refactoring Benefits

The refactored version provides:

  • Modular Code: Separated concerns into config, data, and plotting modules
  • Reusability: Functions can be imported and reused across notebooks
  • Maintainability: Easier to modify settings and add new visualizations
  • Consistency: Standardized plotting styles and data processing

Adding New Charts

  1. Use existing data processing functions from data_utils.py
  2. Apply consistent styling with plot_utils.py helper functions
  3. Follow naming conventions for output files

License

MIT License - see LICENSE file for details.

Contributing

  1. Follow the modular structure when adding features
  2. Update configuration in config.py for new distributions or settings
  3. Add new utility functions to appropriate modules
  4. Test with the refactored notebook before committing

Data Notes

  • EPEL data represents systems checking for package updates
  • Some dates are excluded due to known data collection issues
  • DockerHub analysis is currently disabled pending data source updates
  • System age categories: <1 week (ephemeral), 1 week-1 month, 1-6 months, >6 months (long-term)