Papers I find interesting, read closely, and rebuild: a plain-language summary, the math reimplemented from scratch in TypeScript, live charts, and — where the idea maps onto a 2D toy — a link to poke at it on the model pages. Same rules as the rest of the site: no ML libraries, every number computed in your browser.
A peer-to-peer electricity DQN improves settlement, but its actions cannot affect the next state—and simple budget-balanced rules still win.
A trace-blind difficulty score reaches AUROC 0.873, exposing why pooled reasoning-probe accuracy can vanish within the same problem.
Tabular foundation models lead 316-equation benchmarks, yet retain an uncertainty floor and ignore the unit structure that makes physical laws simpler.
A leakage-free forecasting study shows that baseline warmup alone can move adaptation's apparent benefit by 18.8 percentage points.
A pressure-mat detector reaches 0.969 patient-held-out AUC, but ROC AUC is undefined for nearly one in five patients.
An EV diffusion model beats a real-vs-real Wasserstein baseline, but the metric can ignore when every current spike occurs.
QLoRA sharply improved financial sentiment classification, but none of 28 return-prediction tests survived robust multiple-testing correction.
An adversarial critic slowed reward hacking under a weaker LLM judge—but unconstrained debate defaulted to a different exploit.
A 892-curve field study shows why fertilizer recommendations must be scored by forgone profit, not prediction error—and why its strongest result is a correction that transfers.
A spectral foundation model looked robust to preprocessing until its parameter-free input normalization matched—and sometimes beat—the encoder.
FraudBench finds 3.7 feasible attacks after filtering, but 2,832.3 when the same constraints guide generation—and cost is still missing.
Sparse-looking attention is not a cache certificate: retained mass, omitted values, future queries, and the task can each reverse the conclusion.
A neural PDE surrogate can stay accurate and solver-like while a trigger selects the wrong viscosity—and its headline success metric can still mislead.
INT4 barely moves aggregate accuracy yet loses up to 12.7 points when an LLM must retrieve the latest of many similar updates.
A semantic cache can report 60% hits while only 1.6% of requests reuse a valid answer; admission quality matters before eviction cleverness.
CJSD separates input shift from mechanism change with two classifiers—but vanishing support overlap can make a tiny score inconclusive.
One pretraining row was learned, later undetectable on its own text, yet left a large weight-space displacement: influence depends on when and how you measure it.
Leaf-value coordinates explain boosted-tree decisions exactly, but conventional recourse validity ignores whether a person can make the changes.
Power Sampling can raise correct-trajectory mass yet break self-consistency by collapsing diverse support onto one wrong path.
SAE ablations look like causal feature tests, but letting each dictionary choose its strongest token can manufacture disagreement.
Krum rejects distant client updates; we rebuild its geometry and show why a defense-aware backdoor aims for the cluster center instead.
A clean Uniswap fee formula looks like implied volatility, but one missing economic leg prevents that interpretation.
Tent adapts a deployed model by minimizing prediction entropy; we rebuild the update and show why an unlabeled batch—not confidence alone—does the work.
PCGrad projects away clashes between task gradients; we rebuild the surgery and expose why sequential projections need randomized task order.
Integrated Gradients satisfies elegant attribution axioms, but every explanation is relative to a baseline; we rebuild the path integral and expose that choice.
CLIP turns text prompts into classifier weights; we rebuild its contrastive loss and show why the available labels change the answer.
Focal loss rescued dense object detection by muting easy negatives; we rebuild its weighting rule and expose the hard-example assumption it makes.
CatBoost turns categories into ordered target averages; we rebuild the encoding and expose both the leakage it prevents and the cold-start cost.
Speculative decoding verifies cheap guesses in parallel; we rebuild its exact sampler and find speed depends on agreement and spare compute.
ViT turns image patches into ordinary Transformer tokens; we rebuild the input pipeline and find scale—not patching alone—made it beat convolutions.
FlashAttention streams softmax tiles instead of storing an attention matrix; we rebuild the exact update and find the arithmetic stays quadratic.
DPO turns preference learning into logistic regression; we rebuild the loss and find every judgment is measured relative to a reference model.
LoRA replaces a dense weight update with two skinny matrices; we rebuild the merge and find the tiny adapter never replaces the frozen base model.
TIGER recommends products by generating semantic tokens; we rebuild its quantizer and find a three-order-of-magnitude arithmetic slip.
Exact match can turn smooth language-model improvement into an apparent breakthrough; we rebuild the metric and expose its hidden resolution limit.
DeepSeek-R1 used group-relative RL to reward better answers; we rebuild GRPO and find that all wins or all losses teach it nothing.
KANs replace weights with learnable splines; we rebuild the edge math and find the accuracy grid also taxes every connection.
Keeping old data stops recursive training error from exploding; we rebuild the proof and find “avoids collapse” does not mean nothing is forgotten.
Mamba replaces attention with a selective recurrence; we rebuild its memory gate and find why the fast algorithm is part of the idea.
Deep Hedging learns option strategies under real trading costs; a seeded reconstruction shows why its most useful action can be doing nothing.
Muon orthogonalizes matrix gradients and reportedly halves LLM training compute; we rebuild it and find raw update size silently depends on matrix shape.
BitNet packs 2B weights into 0.4 GB without losing much benchmark score; we rebuild its ternary math and find the savings begin after training.
A 2023 paper forecasts the future with a single linear layer and beats a stack of specialised Transformers. We rebuild it from scratch — and find the clever-looking decomposition barely matters, while the whole edge rests on one assumption nobody states.
After enough backtests, a Sharpe of 2 is what pure luck looks like. A 2014 paper computes exactly how much luck — and deflates the number back down. We rebuild it from scratch and find the verdict hinges on one figure nobody reports: how many strategies you tried.
A 2015 paper shows a tiny, invisible nudge flips a confident classifier — and argues the cause is that models are too LINEAR, not too complex. We rebuild the fast gradient sign method from scratch and watch the attack strengthen as the input grows.
NGBoost turns gradient boosting into a probability distribution with error bars. We rebuild it from scratch and find the whole thing hinges on one correction the title names and the demos gloss over: the natural gradient.
Adam trains almost every deep model — and its original convergence proof was wrong. A 2018 best paper builds a one-line convex problem where Adam walks to the worst point. We rebuild it from scratch and watch a one-word fix pull it back.
Sparse Mixture-of-Experts got the headline — a trillion parameters, a sliver of compute. We rebuild it from scratch and find the real trick is a one-line auxiliary loss: without it the gate collapses onto a few experts and the rest die.
A trading paper bolts a second model onto the first: keep its direction, learn which of its calls to trust. We rebuild the filter from scratch — and find it only adds value when it sees something the first model never did.
A 2019 paper throws away the hand-tuned momentum rule and trains a network to maximise the Sharpe ratio directly. We rebuild the Sharpe loss from scratch — and find the number it optimises is the one it never has to pay for.
A 2017 paper shows modern nets are confidently wrong — and that dividing the logits by one number fixes it. We rebuild ECE and temperature scaling from scratch, then find the one knob can make calibration worse.
Isolation Forest scores anomalies by how fast random cuts fence a point off — but its axis-parallel cuts smear a coordinate-frame bias across the score map. We rebuild it from scratch, expose the phantom corridors, and watch one tilted-cut change erase them.
A microstructure paper shows the mid-price move over a few seconds is a near-linear function of order-flow imbalance — and integrating the whole book explains ~84% of it. We rebuild OFI from scratch, then watch it forecast almost none of the next move.
A 2016 paper builds a diversified portfolio without ever inverting the covariance matrix — and beats the textbook optimiser out-of-sample. We rebuild it from scratch and find its edge is all estimation error, not a better objective.
A 2024 paper turns return forecasts into intervals with a distribution-free coverage guarantee, then picks portfolios from them. We rebuild it from scratch — and watch the guarantee quietly break the moment the market changes regime.
A NeurIPS 2022 paper explains why gradient-boosted trees still beat neural nets on tabular data. We reproduce its sharpest test from scratch: a random rotation flips the ranking, because a tree lives in its coordinate frame.
A Journal of Finance paper times the market with more parameters than data points — and wins. We rebuild its random-feature ridge from scratch, watch double descent appear, and find where the money actually comes from.
A gradient-boosting paper predicts next-day S&P 500 returns and builds a portfolio. We reproduce its indicators from scratch, chart its results — and read the Kelly number it buries.