Market Eyes Live · Quantitative Research Working Paper No. 01
Classification: Public Open methodology Version 1.0 11 August 2026

Validating a Systematic Equity Rating Engine: Selection, Tail Risk, and Portfolio Construction

An open, reproducible validation of the MELANY engine, with every claim stated alongside its sample, statistic, and limitations.

Abstract

The MELANY equity rating engine's conviction score sorts forward returns monotonically across the United States equity market, its price-inflation flag identifies a five-to-twelve-fold elevation in eight-week drawdown probability, and two of its six rule-based model portfolios exceed their natural benchmark on both return and risk. These results are established under deliberately hostile test conditions: every test is point in time, companies that later delisted remain in the sample, and each figure reproduces from a frozen data snapshot.

On 135,605 monthly observations across a delisted-inclusive United States equity universe (2021 to 2026), forward win rate and median return rise monotonically with the engine's conviction score. The top-minus-bottom conviction quintile six-month win-rate spread is 12.0 percentage points (95% CI 8.7 to 15.1; bootstrap probability of a positive spread, 100%). A price-inflation flag isolates high-conviction names that are 5.1 times more likely to draw down 20% and 11.7 times more likely to draw down 30% within eight weeks; the separation is stable across an out-of-sample split and survives a Monte Carlo with a scrambled-label placebo (p = 0.001). Of six rule-based model portfolios backtested net of costs, two exceed their natural benchmark on both compound return and risk-adjusted return.

The measured effects are modest and relative: the engine sorts better names from worse and flags elevated tail risk, but it does not forecast the return of any individual security, and the portfolio results are hypothetical. We state these limits explicitly throughout, and we retire tests we cannot reproduce.

Keywords: survivorship-free backtest, point-in-time fundamentals, cross-sectional forward returns, drawdown probability, cluster bootstrap, portfolio construction, out-of-sample validation.

§Findings in brief

  1. Selection. Forward win rate and median return rise monotonically with conviction; the top-minus-bottom quintile six-month win-rate spread is +12.0 percentage points (95% CI 8.7 to 15.1; bootstrap probability of a positive spread, 100%).
  2. Tail risk. Flagged high-conviction names are 5.1× more likely to draw down 20% and 11.7× more likely to draw down 30% within eight weeks; stable out of sample, placebo-controlled at p = 0.001, and operating live in the product.
  3. Construction. Two of six rule-based model portfolios exceed their benchmark on both compound return and Sharpe ratio, net of modeled costs, and every portfolio beats a naive momentum basket on risk.
Table 1. Summary of findings. Each row states the principal statistic, the sample it rests on, and its evidence class.
FindingPrincipal statisticSampleClass
Conviction predicts forward returnsQ5−Q1 6-mo win spread +12.0pp [8.7, 15.1]135,605 obs / 658 datesPoint-in-time, delisted-inclusive
Tier ordering (BUY vs SELL)Win 56% vs 44%; median +2.92% vs −7.07%Same panelPoint-in-time, delisted-inclusive
Price-inflation flag, 20% drawdown5.1× (28.3% vs 5.6%)538 names, 2021–2026In- and out-of-sample; Monte Carlo
Price-inflation flag, 30% drawdown11.7× (15.2% vs 1.3%)SamePlacebo p = 0.001
Model portfolios beating benchmark2 of 6 on return and risk2022–2026, net of costsHypothetical, backtested
Forward (live) recordDaily capture and gradingFrom 6 Aug 2026Out-of-sample, accruing

All figures are defined and sourced in the sections below. Confidence intervals are 95% date-level bootstrap intervals over evaluation dates. Portfolio returns are hypothetical and net of a modeled 10 basis points per rebalance.

1The validation program

A backtest is only as honest as its construction. Two failures account for most overstated quantitative claims: look-ahead, where information not available on the decision date leaks into the signal, and survivorship, where companies that failed are quietly absent from the sample. Both bias results in the favorable direction, and both are avoidable. Every result in this paper is constructed to exclude them.

We adopt four standing rules. (i) Point in time. Fundamentals enter a score only after they were publicly filed, with a 90-day filing lag applied to approximate real availability. (ii) Survivorship-free. The universe includes delisted securities, whose returns realize through the delisting. (iii) Pre-committed decision rules. The threshold that would move a model into production is written down before the data is examined, which is the discipline that guards against overfitting a backtest to noise (Bailey, Borwein, López de Prado, and Zhu, 2014).3 (iv) Reproducibility. Each result is produced by a named script over a frozen data snapshot and returns the same numbers on re-run.

We also separate questions by how quickly they can be answered honestly, rather than collapsing them into the single, near-unanswerable question of whether a strategy beats the market. That layered framework is set out in Section 6. The short version: we make claims that converge on a business timescale (does conviction sort returns; does the flag mark tail risk; are the published odds calibrated) and we merely track, without ever quoting as a verdict, the question that does not (long-horizon alpha).

2Data and universe

The engine under test is MELANY/PRISM, the production scorer that assigns every rated United States security a conviction score and a tier. Fundamentals (gross and operating margin, return on equity, leverage) and earnings are drawn from SEC EDGAR filings through a no-look-ahead extractor that selects, for each decision date, the most recent fiscal year whose filing date plus 90 days precedes that date. Valuation is priced against the cross-sectional median. Momentum and risk inputs are computed from split-adjusted daily bars. Components for which no point-in-time value can be reconstructed (analyst sentiment, earnings surprise, catalyst proximity, macro regime) are excluded from the reconstructed composite exactly as the live engine excludes empty components, rather than being imputed.

The equity universe is the full bars-complete set of United States listings, survivors and delisted alike. For the selection study this yields 135,605 monthly observations across 658 distinct evaluation dates, of which 56,868 carry reconstructed point-in-time fundamentals; the remainder are scored on the technical and valuation components the engine would use when fundamentals are thin.

3Study I: Conviction and forward returns

The central question is whether the engine's conviction score carries information about subsequent returns. Each security is re-scored on its own rolling 21-trading-day (monthly) cycle, so evaluation dates are staggered across the calendar rather than aligned to month ends. On each evaluation date we rank that date's scored names into quintiles by composite score and measure the forward win rate (share of positions with a positive return) and median return at three and six months. Headline quintile results are computed on the fundamentals-bearing subset (56,868 observations); the full delisted-inclusive panel serves as the survivorship-free cross-check. Because this universe contains squeeze-and-collapse names whose arithmetic mean returns are meaningless, we report robust statistics: win rate, median, and a winsorized mean.

Inference is by a date-level cluster bootstrap: the top-minus-bottom quintile spread is first computed within each evaluation date having at least ten scored names (658 such dates), and those per-date spreads are then resampled with replacement (2,000 draws) to form the interval. This respects cross-sectional dependence within a date; serial dependence across overlapping forward windows is not separately modeled, which we note as a limitation.

3.1Results

Forward performance rises monotonically with conviction. The lowest quintile is a net loser over six months and the highest a clear winner, on both the probability of a gain and the median gain.

Table 2. Forward performance by conviction quintile, fundamentals-bearing subset, 2021 to 2026.
Conviction quintile3-mo win6-mo win6-mo median6-mo wins. mean
Q1 (lowest)44%45%−4.70%3.39%
Q248%48%−1.43%3.79%
Q349%48%−0.82%3.44%
Q450%51%+0.73%3.54%
Q5 (highest)53%56%+2.59%5.14%
Q5 − Q1 spread8.9pp12.0pp7.29pp1.75pp

Notes. Quintiles are formed within each evaluation date on the fundamentals-bearing subset (56,868 observations; per-quintile counts 10,996 to 11,669 at six months). Spread inference uses the 658 evaluation dates with at least ten scored names. Quintile win rates are displayed rounded to the nearest point; the spread rows are computed on unrounded per-date rates, which is why the six-month spread reads 12.0pp rather than the 11pp implied by the rounded endpoints. Win-rate spreads: 6-month 12.0pp, 95% CI [8.7, 15.1]; 3-month 8.9pp, 95% CI [5.8, 12.1]; bootstrap probability that the spread exceeds zero, 100% at both horizons. Winsorization clips returns below −95% and above +200%. By published tier: BUY-rated names realized a 56% six-month win rate and a +2.92% median, versus 44% and −7.07% for SELL; the fully survivorship-free all-universe cut agrees (BUY 58% win / +2.07% median vs SELL 44% / −6.82%).

50% (no skill) Q1 lowest conviction: 45% six-month win rate 45% Q2: 48% 48% Q3: 48% 48% Q4: 51% 51% Q5 highest conviction: 56% 56% Q1 low Q2 Q3 Q4 Q5 high Bars show six-month win rate as deviation from the 50% no-skill line.
Figure 1. Six-month forward win rate by conviction quintile, plotted as deviation from the 50% no-skill line. The relationship is monotone and crosses into positive territory only in the top two quintiles. Source: PRISM selection audit, 20 June 2026; survivorship-free point-in-time reconstruction, n = 135,605.

3.2Factor attribution

Transparency cuts both ways. Attributing the quintile spread to individual components shows that the multi-factor composite does not beat its single strongest input on this reconstructable subset: the risk-adjusted component alone produces a wider spread than the blend. We report this because it bears directly on how the score should be weighted, and because a validation that only reported flattering cuts would not deserve the reader's trust.

Table 3. Six-month Q5−Q1 win-rate spread when ranking by each component versus the composite.
Ranked bySpread95% CIData coverage
Composite (production)12.0pp[8.7, 15.1]100%
Risk-adjusted16.6pp[13.1, 19.8]100%
Quality9.1pp[6.0, 12.0]100%
Momentum8.2pp[5.0, 11.2]100%
Valuation5.1pp[2.0, 8.4]54%

Notes. All confidence intervals in this table come from the same date-level bootstrap procedure as Table 2; the composite row repeats the primary estimate. The risk-adjusted edge is partly the low-volatility anomaly and a mechanical win-rate advantage for low-variance names (Baker, Bradley, and Wurgler, 2011);4 it should be confirmed on median return, not win rate alone, before any re-weighting. This is the point-in-time reconstructable subset and excludes live sentiment and earnings components.

3.3Limitations

The edge is real but modest and relative: it separates better-ranked from worse-ranked names by roughly ten to twelve points of win rate, not a per-security guarantee. The reconstruction omits two components the live engine uses (analyst sentiment and earnings surprise), which are not point-in-time recoverable and are therefore untested here. The fundamentals-bearing subset is mildly survivor-leaning because delisted issuers rarely resolve to a current filing entity, though the all-universe cut, which is fully survivorship-free, corroborates the tier ordering. The 90-day filing lag is an approximation of exact availability.

4Study II: Price inflation and tail risk

A security can carry a high conviction score and simultaneously accumulate crash risk. This study tests a deliberately narrow claim: that a price-inflation flag on high-conviction names identifies elevated drawdown probability, not lower average returns. The flag fires when at least two of six conditions hold: a stretch z-score at or above 2, price more than 60% above the 200-day moving average, parabolic acceleration, a 14-day RSI at or above 78, a volume climax, and a three-month advance of 60% or more.

Outcomes are measured over 40 trading days (about eight weeks) on 538 names from 2021 to 2026, with an in-sample and holdout split. We record the probability of a peak-to-trough drawdown of 20% and of 30% for flagged versus unflagged high-conviction names.

20% drawdown, flagged: 28.3% 28.3% 20% drawdown, not flagged: 5.6% 5.6% 30% drawdown, flagged: 15.2% 15.2% 30% drawdown, not flagged: 1.3% 1.3% Drawdown ≥ 20% Drawdown ≥ 30% 5.1× more likely 11.7× more likely Flagged (price-inflated) High conviction, not flagged
Figure 2. Probability of a large drawdown within eight weeks, flagged versus unflagged high-conviction names. In-sample and holdout probabilities agree to within 1.5 points (20% drawdown: 27.8% in-sample, 29.3% holdout). Source: top-risk-signals backtest and Monte Carlo, 28 July 2026; 538 names, 2021 to 2026.
Flagged high-conviction names were 5.1 times more likely to fall 20% (28.3% versus 5.6%) and 11.7 times more likely to fall 30% (15.2% versus 1.3%) within eight weeks. A Monte Carlo across all five years, with a scrambled-label placebo and symbol-level bootstrap, placed the effect at p = 0.001, with the separation positive in every individual year and robust to removing any single security (worst case, +30 points).

Crucially, the flagged names' average forward returns were the same or higher than the unflagged group. This is the pattern documented by Greenwood, Shleifer, and You (2019): sharp run-ups forecast crash probability, not low mean returns.1 The correct product response is therefore to size positions down and raise the alert on flagged names, which is how the signal operates live as the Overheated state, rather than to sell or to lower the rating.

Before deployment, the flag was re-validated under a Monte Carlo across all five sample years with a scrambled-label placebo and symbol-level bootstrap (95% interval on the flagged-minus-unflagged separation, +26 to +36 points), and the production condition set measured a wider separation than the research definition reported above (36.9% versus 5.7% at the 20% threshold). The figures in this section are the research definition's, which is the more conservative of the two. The limitations are the mirror image of the claim: the flag speaks to drawdown probability, not expected return; it is specific to the high-conviction subset; and it is measured at an eight-week horizon.

5Study III: Rule-based model portfolios

The final study asks whether the engine's outputs support disciplined portfolio construction. Six model portfolios are built by rule, rebalanced monthly on point-in-time eligibility with delisted names retained, and charged a modeled 10 basis points per rebalance. The construction imposes position caps, drift bands, and a regime brake. The primary window is 2022 to 2026 on a universe of roughly 1,100 names; a stress window extends to 2008 to 2026 and spans the 2008 financial crisis, the 2020 pandemic drawdown, and the 2022 bear market. All portfolio results are hypothetical and backtested.

The evidence that construction, rather than luck, is doing the work is that every portfolio, all six, clears a naive top-six momentum basket (43.1% compound annual return, but at 59% volatility with a 57.2% maximum drawdown) on both Sharpe ratio and drawdown, in the primary window and again in the stress window. Two portfolios also exceed their natural benchmark outright.

Table 4. Model-portfolio results versus benchmark, primary window, net of costs. Hypothetical and backtested.
Model portfolioCAGRSharpeMax DDBenchmark (CAGR / Sharpe)Read
MELG Growth40.7%1.42−30.3%QQQ · 24.5% / 1.16Exceeds on return and Sharpe
MELD Steady12.5%0.93−12.5%SCHD · 10.6% / 0.75Exceeds on all three
MELX Core22.7%0.89−28.2%SPY · 16.7% / 1.01Higher return, higher risk
MELB All-Weather13.8%0.99−17.3%SPY · 16.7% / 1.01Lower-volatility profile
MELC AI & Chips41.8%1.19−44.1%SMH · 51.3% / 1.47Trails a concentrated index

Notes. Sharpe and maximum drawdown are net of a modeled 10 basis points per rebalance. The benchmark column reports the index compound return and Sharpe over the same window. MELC is reported deliberately: a diversified twelve-name book does not out-return a pure semiconductor index inside a semiconductor bull market, and omitting the case would misrepresent the method. The lineup's sixth fund, the speculative sleeve (MELS), is excluded from this tearsheet because its historical statistic derives from a momentum-proxy study of the sleeve's universe rather than the admission strategy as deployed, and we withhold numbers we cannot attribute to the deployed strategy; its rule-based construction nonetheless clears the naive baseline on both Sharpe ratio and maximum drawdown in both windows. Figures reflect the corrected cap-enforcement build (v2, 8 August 2026); an earlier build overstated the two most concentrated portfolios and was superseded. Deep-window (2008 to 2026) results are retained on file.

Limitations. These are hypothetical returns, not a live track record, and are subject to the standard limits of backtested performance. A concentrated theme index can and does out-return a diversified book within its own boom, as the MELC row shows. The value on offer is risk-controlled construction, evidenced by drawdowns materially shallower than the naive baseline and, in the growth sleeve, shallower than the index across the stress window.

6Model governance

We treat "does the model beat the market" as close to unanswerable on a business timescale, because separating skill from luck for volatile securities requires a decade or more of data. Rather than treat each six-week window as a verdict, we decompose evaluation into five layers, each with its own instrument and its own honest clock.

6.1The five-layer framework

Table 5. Evaluation layers and the decision each may drive.
LayerQuestionConverges inMay drive
0  MoneyDid a follower make money, in dollars, as prescribed?ImmediateHonesty of any published figure
1  PromisesDoes the product do what the card says (stops, entries, freshness)?DaysShip or fix, always
2  CalibrationWhen it says 60% odds, is the realized frequency near 60%?WeeksCopy, confidence display
3  OrderingDoes BUY beat HOLD beat SELL within the universe?QuartersTier thresholds
4  AlphaDoes it beat the market outright, over a decade?RarelyNothing. Tracked, never a verdict

Notes. The governing rule is the last row: long-horizon alpha is recorded and published honestly with its window stated, but it never drives a product decision, because on the relevant timescale that estimate is dominated by noise.

6.2Rejected modifications

A model that only ever adds features is a model that never says no. Of more than fifty formal studies conducted under the standards in Section 1, the majority of proposed scoring changes were declined. A representative sample follows; each was measured, failed its pre-committed bar, and was not shipped.

Table 6. Proposed modifications tested and declined.
Proposed changeReason declinedVerdict
Faster momentum average (20-day vs 50-day)Edge vanished at 5× sample; added churnRejected
Piotroski F-score into quality (Piotroski, 2000)2Regime-dependent, redundant with existing quality inputsRejected
Sector-relative quality gradingSlightly worse than the absolute scale on forward returnsRejected
Additional relative-strength wiringRedundant with existing momentum; diluted the compositeRejected
Factor re-weighting toward momentumImproved win rate but not median return, out of sampleRejected
Early sector-rotation timing modelTwo attempts on ten years of data; both decayed out of sampleRejected
"Wait for the pullback" entry adviceDid not beat entering at market; filled only 37% of the timeReversed
Automatic take-profit and stop overlaysExisting exits already cap the tail; overlay cost returnRejected

Notes. Pre-committed decision rules mean the pass threshold is fixed before results are seen, so a study that reverses under a more conservative test (for example, month-clustered rather than event-weighted significance) is caught rather than shipped.

7Forward validation

Backtests describe the past. Their honest counterpart is a forward record that cannot be curve-fit because it has not yet occurred. Since 6 August 2026, every rating the engine publishes, including its conviction score, tier, suggested entry, and risk flags, is captured daily and graded against realized forward price action on a fixed schedule.

This record is deliberately young, and we treat it accordingly. Nothing from it is quoted as a verdict until the sample and the number of distinct market regimes are sufficient to support the claim, with pre-committed sample floors (a minimum count of observations and of distinct securities per reported cell). The value of the record is that it accrues in the open and answers, over time, the one fair question every backtest invites: what happens live.

8Scope of claims

Each result above is stated at its exact strength, no more and no less. For the avoidance of doubt: the selection edge is a relative sort of the universe, roughly ten to twelve points of win rate between top and bottom conviction, and not a per-security forecast. The price-inflation flag forecasts drawdown probability, not expected return. The model-portfolio figures are hypothetical and backtested, net of a modeled cost only. Long-horizon alpha is tracked in the live record with its window stated, and is not asserted as a settled verdict for the reasons in Section 6.

The same precision governs quality control. Any internal study that fails reproduction is retired: one early validation was removed after our own audit found a graded outcome had leaked into a model input, and it contributes to none of the figures in this paper. The results that remain are the ones that survived that filter.

9Reproducibility

Each result above is generated by a named script executed over a frozen snapshot of its input data, and returns identical numbers on re-run. Snapshots are retained so that a result can be regenerated with the exact state present at test time rather than re-fetched from live sources. Definitions used throughout:

Point in time
A signal on date t uses only data with a filing or observation date at or before t, with a 90-day fundamentals filing lag.
Survivorship-free
The universe retains delisted securities, whose returns realize through the delisting rather than being dropped.
Win rate
Share of positions with a strictly positive forward return over the stated horizon.
Winsorized mean
Mean after clipping each return to the fixed band of −95% to +200%, limiting the influence of squeeze-and-collapse outliers.
Date-level cluster bootstrap
The statistic is computed within each evaluation date, and dates are then resampled with replacement to form the interval. Dependence within a date is respected; dependence across overlapping forward windows is not separately modeled.
Maximum drawdown
Largest peak-to-trough decline in portfolio value over the window.

10Conclusion

Under assumptions chosen to make success difficult, with look-ahead excluded, failures retained, decision bars fixed in advance, and every figure reproducible from a frozen snapshot, the engine's central claims hold. Conviction sorts forward returns monotonically across the market. The price-inflation flag isolates a five-to-twelve-fold elevation in tail risk and does so identically out of sample. Rule-based construction converts those signals into portfolios that, in two of six cases, beat their benchmark on both return and risk, and in all six cases beat naive construction on risk.

Since 6 August 2026, the same engine has been grading itself in public, daily, against realized prices. Every week that record grows, the distance between what we measured and what we can demonstrate live gets shorter. That is the standard we invite the reader, and the skeptic, to hold us to.

RReferences

  1. Greenwood, R., Shleifer, A., and You, Y. (2019). "Bubbles for Fama." Journal of Financial Economics, 131(1), 20–43.
  2. Piotroski, J. D. (2000). "Value Investing: The Use of Historical Financial Statement Information to Separate Winners from Losers." Journal of Accounting Research, 38 (Supplement), 1–41.
  3. Bailey, D. H., Borwein, J. M., López de Prado, M., and Zhu, Q. J. (2014). "Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance." Notices of the American Mathematical Society, 61(5), 458–471.
  4. Baker, M., Bradley, B., and Wurgler, J. (2011). "Benchmarks as Limits to Arbitrage: Understanding the Low-Volatility Anomaly." Financial Analysts Journal, 67(1), 40–54.

FFrequently asked questions

Does the MELANY stock rating engine predict returns?

In a survivorship-free test of 135,605 monthly observations (2021 to 2026), forward win rate and median return rose monotonically with MELANY's conviction score. The top conviction quintile beat the bottom by 12.0 percentage points of six-month win rate (95% CI 8.7 to 15.1). The edge is a relative ranking of the market, not a forecast for any single stock.

Can MELANY warn about a crash before it happens?

Its price-inflation flag identifies high-conviction stocks that went on to suffer a 20% drawdown within eight weeks 5.1 times more often than unflagged peers, and a 30% drawdown 11.7 times more often, stable out of sample and placebo-controlled at p = 0.001. It forecasts drawdown probability, not returns, so the product responds by reducing position size, not by selling.

Are the model portfolio backtests survivorship-biased?

No. Portfolios are rebalanced monthly on point-in-time eligibility with companies that later delisted retained in the universe, net of modeled costs. Results are hypothetical and backtested; two of six portfolios exceeded their benchmark on both return and Sharpe ratio, and the paper reports the one that trails its index.

Does Market Eyes Live have a live track record?

Since August 6, 2026, every published rating is captured daily and graded against subsequent market moves on a public record that started before results existed. Raw counts are published at marketeyeslive.com/api/validation-status, and nothing from the live record is quoted as a verdict until pre-committed sample floors are met.

CHow to cite

Market Eyes Live Quantitative Research (2026). “Validating a Systematic Equity Rating Engine: Selection, Tail Risk, and Portfolio Construction.” Market Eyes Live Working Paper No. 01, August 2026. https://marketeyeslive.com/melany-validation-study.html

Disclosures

Hypothetical and backtested performance has inherent limitations, is prepared with the benefit of hindsight, and does not represent actual trading or the effect of material economic and market factors on decision-making. Past performance is not indicative of future results. Model-portfolio figures are net of a modeled transaction cost only and do not reflect advisory fees, taxes, slippage, or capacity constraints.

This document is for informational purposes and describes research methodology and results. It is not investment advice, a research recommendation, or an offer or solicitation to buy or sell any security or to adopt any strategy. Statistics are point-in-time reconstructions produced from retained data snapshots; methods and snapshots are available for inspection, and replication is invited.

Market Eyes Live · Quantitative Research Working Paper No. 01 · v1.0 · 11 August 2026