Oregon School Data Explorer

Spending and Class Size vs Performance: Findings Memo

Published analysis report from Evidence Lab artifacts.

Spending and Class Size vs Performance: Findings Memo
Regenerated: 2026-09-08

Scope
- Data: 2024-2025 processed English, Math, Science school rows.
- Outcome: Percent Proficient, reconstructed as Number Proficient / Scored Performance Denominator.
- Predictors of interest: Overall spending per student, Classroom spending per student, Median class size.
- Context controls (baseline continuity spec): Per-capita income, Adult BA+ rate, Regular attendance.
- Companion context spec: Students Experiencing Poverty, with censored AAG disclosures represented by documented interval midpoints and their raw tokens/bounds preserved upstream.
- Weighting: students with reported Level 1-4 results.
- Grade-record attendance must carry exact grade-population provenance; censored or otherwise nonnumeric attendance is omitted without substitution.

Methods used
1) Weighted bivariate checks (correlation and simple slope).
2) Weighted multivariable OLS-style checks with controls.
3) Weighted ridge regression with five-fold school-clustered CV (to stabilize collinear spending variables without splitting one school's grade records across folds).
4) District fixed-effects ridge (within-district demeaning) with school-clustered CV and permutation importance.

Executive summary
- Your chart-level impression is largely confirmed:
  class size shows little robust predictive contribution once SES and attendance are included.
- Spending shows a detectable but inconsistent contribution and is highly sensitive to context specification and model frame.
- Education and attendance remain dominant in pooled models and many fixed-effects models, but not every poverty-controlled fixed-effects result.
- Income is generally intermediate: weaker than education and attendance, often similar to or above spending.
- The poverty-aware fixed-effects sensitivity is a qualification: the combined spending contribution is notably larger for Math and Science than in the pooled poverty models, so a single project-wide rank ordering is too strong.

Key findings by factor
1) Median class size
- Bivariate association can appear positive.
- After controls and regularization, standardized effects are near zero in all subjects.
- Permutation importance is near zero in most settings.
- Interpretation: class size is not a strong standalone predictor in these school-level cross-sections.

2) Spending (overall/classroom)
- The two spending variables are highly collinear (weighted corr ~0.85 to 0.88).
- Ridge regularization was necessary to reduce coefficient instability.
- Combined spending contribution (permutation drop in CV R2):
  - Per-capita-income spec, non-FE ridge: English 0.047, Math 0.012, Science 0.047.
  - Per-capita-income spec, FE ridge: English 0.404, Math 0.080, Science 0.386.
  - Poverty spec, non-FE ridge: English 0.003, Math 0.012, Science 0.005.
  - Poverty spec, FE ridge: English 0.006, Math 0.219, Science 0.010.
- Interpretation: spending can make a secondary predictive contribution, but its strength is uneven across subjects, context controls, and model frames.
  The English/Science income-spec FE results and the Math poverty-spec FE result should be treated as sensitivity findings, not stable effects, until replicated with lagged spending and alternative fixed-effects specifications. Their permutation magnitudes changed substantially when CV was corrected to keep each school's grade records in one fold, which is itself evidence that these estimates are not stable headline effects.

3) Education, attendance, income
- Education and attendance produce the largest predictive contribution in the pooled models and most income-spec fixed-effects results.
- Math FE ridge particularly emphasizes attendance.
- Income remains meaningful but below education/attendance in most configurations.
- Under the poverty-spec fixed-effects lens, combined spending exceeds BA+ in Math and Science, but remains below attendance in Science. This exception is why the findings are described as model-sensitive rather than placed in a single ordering.

How to interpret the apparent contradictions
- Sign flips and coefficient changes across models are expected under strong collinearity
  (especially between overall and classroom spending).
- Permutation importance is more stable for relative contribution than raw coefficient signs.
- FE models answer a different question (within-district differences) than pooled models
  (between + within combined), so effect sizes are not directly interchangeable.

What this means for the project
- It is reasonable to say:
  "At this stage, class size remains weak, while spending is model-sensitive: usually secondary in pooled statewide models but more prominent in several fixed-effects specifications."
- It is not yet reasonable to say:
  "Spending has no effect."
  This data structure (single-year spending, high collinearity, cross-sectional school aggregates) limits causal interpretation.

Recommended next analyses (highest value first)
1) Lag-aligned resource models
- Match outcomes to prior-year or multi-year averaged spending/class-size where possible.
- Rationale: test outcomes may respond to resource conditions with delay.

2) FE + regularized nonlinear checks
- Add spline terms for spending and attendance under ridge/elastic-net.
- Rationale: linear effects may understate threshold or diminishing-return structure.

3) Spending decomposition
- Replace or augment the available spending measures with more policy-proximal components
  (instructional share, staffing mix, support services, etc., if available).
- Rationale: total spending may mask the relevant channel.

4) Robustness by school level (elementary/middle/high)
- Run the same ridge/FE stack by school level.
- Rationale: resource-performance coupling may differ by grade span.

Files produced
- Technical model report:
  docs/spending_class_size_ridge_report.txt
- Companion OLS-style report:
  docs/spending_class_size_effects_report.txt
- Scripts:
  scripts/report_spending_classsize_ridge.py
  scripts/report_spending_classsize_effects.py
Launch dashboard