Methodology方法论
How the Factor Lab, Thematic Baskets and Seasonality pages are built — every formula, lookback and honest limit, on FREE public data (SEC EDGAR + a ~3-year S&P-1500 price cache + the Ken French library). Read this first; it is what makes the numbers interpretable rather than impressive.因子实验室、主题篮子与季节性页面的构建方式 — 每个公式、回看窗口与诚实的局限,全部基于免费公开数据(SEC EDGAR + 约 3 年标普 1500 价格缓存 + Ken French 库)。先读这里。
Overview & honesty概述与诚实声明
This suite computes the cross-sectional equity factors quant desks actually trade, builds them into point-in-time portfolios, and stress-tests whether the ranks predict forward returns. It runs entirely on free data, which sets a hard ceiling we state everywhere rather than paper over:
- Fundamentals are ANNUAL. SEC EDGAR's free XBRL "frames" give one clean cross-section per fiscal year, so fundamental factor ranks (value, quality, etc.) refresh roughly once a year per name. Most month-to-month portfolio turnover comes from the price-derived legs (low-volatility, low-beta).
- History is ~3 years. The daily price cache spans roughly the last three years, so every Sharpe ships with a block-bootstrap confidence interval (which mostly straddles zero) and long-run behaviour is deferred to the Ken French seasonality page.
- Descriptive, not alpha. On a leak-free point-in-time test, few factors clear the multiple-testing bar (see the IC scorecard). These are ranks and context, not a backtested trading system.
Data sources数据来源
| Dataset | Source (free) | Use |
|---|---|---|
| Fundamentals | SEC EDGAR XBRL frames API (keyless) | value / quality / profitability / investment / payout legs, FY2009→latest, stamped point-in-time |
| Prices | Stored S&P-1500 daily close caches (large + mid + small) | returns, volatility, beta, the full daily portfolio/basket series |
| Benchmark | SPY daily close | the cap-weighted market line in every chart + the "vs S&P" relative |
| Seasonality | Ken French Data Library (Dartmouth), monthly factor returns | deep-history (1926/1963+) seasonal climate per style |
| Reference ETFs | Stored ETF daily closes (sector SPDRs, SMH, factor ETFs…) | basket theme cross-check + (planned) factor-vs-ETF validation |
| Short interest / insiders | FINRA bi-monthly / SEC Form-4 (optional) | a standalone short-interest leg + an insider-conviction panel |
Universe & point-in-time股票域与时点
The universe is the S&P 1500 (large + mid + small cap) — the broadest set with free, reliable fundamentals. At a rebalance date t, a fiscal-year report is visible only if its filing became public on or before t (we apply a conservative reporting lag), so no factor ever sees a not-yet-filed statement. Prices are likewise truncated to t. The honest limit: the cache carries today's constituents, so the series is mildly survivorship-biased upward — stated, not hidden.
Factor definitions因子定义
Each factor is a winsorized cross-sectional z-score (higher = more attractive). Composites are the equal-weight mean of available legs.
Value
Mean z of four yields: earnings/price (TTM net income ÷ market cap), book/price (common equity ÷ market cap), sales/price (revenue ÷ market cap), and cash-flow yield (operating cash flow ÷ market cap).
Profitability
Gross profitability (Novy-Marx): gross profit ÷ total assets.
Quality
Mean z of high ROE (net income ÷ equity), low accruals (cash-backed earnings: −(net income − operating cash flow) ÷ assets), and low leverage (−long-term debt ÷ assets) — an AQR Quality-minus-Junk reading.
Investment
Conservative investment (Fama-French CMA): −asset growth (firms that grow the balance sheet aggressively tend to underperform).
Payout
Net shareholder yield: (dividends + buybacks) ÷ market cap.
Low volatility
−trailing 252-day idiosyncratic volatility (the low-vol anomaly).
Low beta
Betting-Against-Beta (Frazzini-Pedersen): −rolling 252-day market beta vs SPY.
Standalone legs also computed but kept out of the composite: accruals (a quality sub-leg) and short interest (a current FINRA snapshot, so excluded from any historical series to avoid leakage).
Cross-sectional z-scores横截面 z 分数
At each rebalance, raw factor metrics are winsorized (clipped at ±3σ to tame outliers and infinities), then z-scored across the point-in-time universe. The composite is the equal-weight mean of the legs a name has (minimum three). A Löwdin-decorrelated composite is also computed so crowding/overlap can be measured (see the IC scorecard's collinearity read).
Portfolio construction组合构建
From the ranks, two series per factor are built and held buy-and-hold between month-end rebalances:
- Long-only — the top quintile, cap-weighted with a 5% single-name cap (excess redistributed pro-rata, so it can't degenerate into a few mega-caps). This carries with the market; compare it to the SPY line.
- Q5−Q1 spread — top quintile minus bottom quintile, equal-weighted — the academic factor-premium diagnostic. It is a paper long/short (not a shortable book) and isolates the premium.
Stats per series: annualized Sharpe with a block-bootstrap 95% CI, max drawdown, and (across the screened factors) a deflated-Sharpe haircut. A green Sharpe whose CI straddles zero is shown as exactly that.
Returns, horizons & σ回报、跨度与 σ
Returns are total returns (split- and dividend-adjusted prices). Horizons: 1d / 5d / 20d / 60d / MTD / YTD, each rebased to the start of the window. The σ ("sigma") mode z-scores the current h-day move against the prior 252 overlapping h-day moves of that same series — "is this move unusual for THIS series?", with |z| ≥ 2 the headline flag. σ is only defined for the fixed-window horizons (1d/5d/20d/60d), not MTD/YTD. Because the windows overlap, σ answers "unusual?", not "statistically significant?".
The IC scorecard — our rigor edgeIC 记分卡 — 我们的严谨之处
This is the honest test most factor dashboards (FactorWatch included) don't show. At each quarter-end we rebuild every factor as it was knowable then (no look-ahead, no survivorship-recovery) and rank-correlate the factor score with the next 63-day return — the Information Coefficient. We report:
- mean IC and IC-IR (IC ÷ its volatility, annualized) — the predictive edge and its consistency;
- tHAC — a Newey-West t-stat (overlapping forward windows autocorrelate the IC series, so the naive t-stat overstates significance);
- qFDR — a Benjamini-Hochberg false-discovery-rate control across the whole factor panel (screening many factors inflates false positives);
- collinearity — pairwise factor correlation + VIF, with the Löwdin-decorrelated composite to show how much overlap is real.
Current read over 2011-03-31..2025-12-31 (60 quarterly rebalances, ~1154 names): few factors clear BH-FDR(10%), and the point-in-time fix made the spreads honestly weaker than a naive latest-data backtest — that gap is the look-ahead/survivorship bias being removed. Judge every number against zero, and against n≈60.
Sector monitor行业监测
Each of the 11 GICS sectors is built into a cap-weighted daily series, expressed relative to the S&P at each horizon, and colored by how unusual that relative move is vs the sector's own trailing-year distribution (the same σ z-score). |z| ≥ 2 flags an outsized rotation. It is a relational view — sector vs market, normalized by the sector's own history — not an absolute-return ranking.
Thematic baskets主题篮子
Curated thematic stock baskets, equal-weighted, with point-in-time dated membership (a member counts only within [added, removed)), rebalanced at month-ends and on any dated change. Each basket carries a written thesis and each member a one-line rationale; the dated changelog is part of the product. Performance is shown raw, relative to the S&P, and in σ; where a listed ETF approximates the theme the basket is cross-checked against it both absolute and market-adjusted (basket−S&P vs ETF−S&P, which strips shared market beta).
Hindsight caveat: membership was curated knowing the period, so the full ~3y series is a backtest of the as-of-creation membership — descriptive structure, not an out-of-sample track and not a buy list (the displacement-style baskets in particular are monitoring lists / hedge candidates).
Seasonality (Ken French)季节性 (Ken French)
Deep-history seasonality uses the Ken French Data Library's monthly factor returns (Momentum 1927+, Value/Size/Profitability/Investment 1963+) — academic total-US-market long/short portfolios, decades longer than our price cache. Per calendar month we report mean, median and hit-rate (% of months positive), over both full history and a trailing-30-year window.
These are a definition mismatch vs our S&P-1500 z-score quintiles — treat them as the long-run climate of a style, not a forecast of our daily factor leadership. Seasonality is deliberately kept out of every calibrated score.
Validation验证
The primary validation is the leak-free IC scorecard above (forward-IC, Newey-West, FDR, bootstrap CIs) — a stricter bar than correlation-to-an-index. A complementary leg cross-checks each factor's long-only series against its corresponding factor ETF (daily correlation, absolute and relative-to-SPY); with no published-index license we validate against free ETFs rather than vendor indices. We do not claim parity with paid multi-decade factor indices.
Limitations & disclaimer局限与免责声明
- Annual (not quarterly) fundamentals; ~3 years (not multi-decade) of prices; mild survivorship bias; ETF-only validation.
- σ uses overlapping windows (autocorrelated); the IC sample is small (n≈60); baskets are hindsight-curated.
- Everything here is read-only analytics for informational purposes only — not investment advice, provided as-is with no guarantee of accuracy, timeliness or completeness.