🔬 Data Integrity & Standards
We believe in full algorithmic transparency. Here is how our data is harvested, scrubbed, normalized, and benchmarked.
1. Multi-Source Aggregation
Single-source salary reports are notoriously biased — employee self-reported reviews skew high or low depending on sentiment, while job boards only capture actively posted rates. StatsForSkills continually synthesizes live vacancy listings and compensation benchmarks across multiple tier-one hiring networks:
Adzuna Live APIs
LinkedIn Jobs
Indeed Market Index
Glassdoor Benchmarks
2. Percentile Modeling (P10, P50, P90)
Averages are easily distorted by extreme outliers (e.g. executive contractor rates or unpaid internships). We use robust statistical percentiles:
-
P10:
10th Percentile — Entry-level or junior threshold compensation in that regional market.
-
P50:
Median Salary — The true market midpoint where 50% of verified postings pay higher and 50% pay lower.
-
P90:
90th Percentile — Senior, staff, or lead compensation packages for top performers.
3. Skill Frequency & 30-Day Velocity
Every month, our skill extraction pipeline analyzes live vacancy descriptions for each job title. We compute:
• Frequency %: Percentage of active postings for that role that explicitly require a specific technology or framework.
• 30-Day Trend Velocity: Compares demand changes between the active 30-day window and the preceding 30–60 day baseline to flag rising (+X% ↗) or cooling (-X% ↘) tools.
4. Sample Size Integrity & Confidence Tiers
We never present unverified guesses as absolute facts. Every salary headline carries an explicit confidence indicator based on live posting volume:
🟢 High Confidence
30+ live verified postings with robust percentile distributions.
🔵 Moderate Sample
10–29 postings sampled from active regional vacancies.
🟡 Directional Estimate
Low-volume or emerging roles utilizing cross-region compensation modeling.