Xgboost Vs Random Forest

  Random forest XGBoost
How trees combine Parallel average Sequential sum
Individual trees Deep: low bias, high variance Shallow: high bias, low variance
Mainly reduces Variance Bias
More trees Never hurts; converges Eventually overfits, so needs early stopping
Tuning Little; defaults work Substantial: learning rate, depth, rounds, \(\lambda\), \(\gamma\), subsampling
Loss function Fixed split criterion (squared error, Gini, entropy) Any twice-differentiable loss: ranking, quantile, robust
Missing values Depends on implementation Native: learns a default direction at each split
Validation needed Out-of-bag estimate, which is inflated under dependence A validation set for early stopping

When each wins

Random forest tends to win when:

  • Signal-to-noise is low. Boosting fits residuals, and when the signal is tiny the residuals are mostly noise, so boosting chases it. Averaging is more forgiving. That’s the reasoning behind your course’s line “in financial applications bagging is generally preferable to boosting.”
  • Validation data is scarce or leaky. Boosting needs a validation set to decide when to stop, and choosing the stopping point on validation data is selection on that data. In finance, clean validation data is precisely the scarce resource.
  • Labels are noisy. Boosting concentrates effort on the hardest cases, and mislabelled points are hard cases.
  • You need a robust baseline quickly, with no tuning budget.

XGBoost tends to win when:

  • There’s a lot of learnable structure, moderate noise and enough data. That describes most tabular benchmarks, where tuned boosting is usually the most accurate option.
  • The structure is many small effects plus low-order interactions. Sums of shallow trees capture that efficiently; a forest needs deep trees.
  • You need a specific loss: a ranking objective for cross-sectional signals, quantile loss, or a robust loss for heavy-tailed targets.
  • Features have missing values, handled natively.

One qualification on financial data: “bagging over boosting” is López de Prado’s position, not consensus. My understanding is that heavily regularised gradient boosting is common in quant practice: shallow trees, small learning rates, subsampling, and early stopping on properly purged folds.