Slower, calmer synthetic speech without slowing the words: render each sentence once at natural speed, then lengthen only the silences the model already made. Every register hits its words-per-second target at speed 1.00.
A transformer inference engine in plain C with no libraries underneath, small enough to run a real chat model offline on a phone in under 90 KB of WebAssembly. Along the way: why a smaller quantization error can make the model worse.
A pocket 60.5 GHz radar that measures movement in micrometres by reading the phase of the returning wave, enough to count a heartbeat through still air or recover sound from a surface it never touches.
A research platform for systematic trading in which Claude designs strategies but never writes code. Each strategy is a frozen, hashed artifact run by a deterministic engine, so the results are reproducible even though the author is not, and every idea has to clear an out-of-sample gate that tightens as more are tried.
A typed graph for modelling an emotional history honestly. Every claim has to declare how it is known, hypotheses are never promoted to facts, and contradictions stay open instead of being resolved.
A Claude Code skill that holds the assistant in one persona for a whole session while the code, comments and commits stay clean. One Markdown file, open source.
A ghostwriting studio for ministry authors where the AI editor cannot invent: every margin note must quote the manuscript verbatim, and a failed check goes back to the model so it corrects itself in the same loop.
Most of my research asks one question: how do you know an automated system is right? My answer is to measure it where the truth is already known: satellite sites certified not to have changed, proof claims that either were or were not checked, a trading gate that has to reject strategies which only looked good in-sample. The first two are the main papers. The rest are shorter: a technical note and system reports on the tools behind that work, a preprint on how many projects to start, and one essay.
Terence Tao predicted that AI-generated proofs would accumulate faster than they could be verified. This measures it. Every public proof claim on erdosproblems.com from July to September 2026, 251 of them and 97% made with AI, was traced to see whether anyone ever checked it. Of the 224 on problems still open, 48% drew any comment, 19% reached any verdict and 4% were accepted by the site's moderators. Verdicts arrive within a day or not at all: after 60 days, 82% of claims still have none. Attaching a Lean proof made no difference, 19 people delivered every verdict, and one complete, formally verified proof was dismissed with a mistaken "already solved" that nobody corrected.
You cannot label the earth, so nobody measures how often a satellite change detector invents change. This runs the pipeline over ground another community already certified as invariant, where every detection is a false alarm by construction. The field has not done this because it cannot: on an empty ground truth, F1 and IoU are undefined for a perfect result and identically zero for every imperfect one, so they cannot tell one false positive from five hundred thousand. On the six CEOS-endorsed desert sites, a threshold chosen from the data reports a median 21.41% of certified-stable ground as changed, against 0.04% with an absolute floor. Otsu calls 39% of an empty scene changed, eases only to 27% as a real signal strengthens, then falls 26.97 points across a single 0.25-sigma step, so behaviour on scenes containing change predicts nothing about scenes without it. Seven defects surfaced, five invisible in the outputs, including a pipeline that reports thirty times less change once a large real structure appears.
The doppelgänger of folklore terrifies because it is a copy of you that is not you. But the body is a pattern rebuilt from new matter, memory is re-saved each time it is opened, and everyone who knows you carries a version of you you will never meet. From Goethe's consoling double and the brain's misfiled selves to Yoruba twin figures and a chatbot built from a dead man's messages, an essay on why there was never an original, and what the copies owe each other.
A synthesized netlist is pure logic, but a layout a foundry could build is logic plus the machinery that lets it survive fabrication: antenna diodes, well taps, timing buffers, fill. Taking one signed int8 multiply-accumulate through the open Sky130 flow at lane counts from 1 to 64, I decompose every placed layout and watch the arithmetic share fall from 59 percent of the cells to 48, crossing below half, because the overhead outgrows the logic and its pieces scale against wiring and die area rather than the gates.
An attempt needs two separate efforts: one to make it work and one to put it in front of anyone. The optimal number of simultaneous attempts depends on their sum, but the heuristic builders actually use depends on build cost alone. Agent assistance collapsed build cost and left shipping cost untouched, so the two prescriptions have diverged. Simulation plus a 48-repository archive where every stalled project is stalled at a step building cannot resolve.
When a language model proposes trading strategies, the trial count that selection-bias corrections depend on is not the number of proposals but the effective number of independent ones. Borrowing the effective-number-of-tests estimator from statistical genetics gives a principled M_eff, and on the live ÀYÒ pipeline the size of the correction turns out to be a property of the prompting protocol: diverse angles yield near-independent proposals that need no correction, while drawing repeatedly from a single narrow angle collapses fourteen proposals to three effective trials and inflates the acceptance bar enough to reject genuine strategies.
A validation gate built to be un-foolable by construction: strategies arrive as frozen, content-hashed data artifacts, as an LLM would propose them, and a deterministic verifier trades only what survives walk-forward out-of-sample testing, a trials-adjusted acceptance bar, and robustness sweeps. Tested here on scripted strategies, on real markets it turned an in-sample Sharpe of 3.07 into a −4.90 rejection and unmasked a +0.13 out-of-sample "edge" as a 0-of-12 fluke — deploying nothing, which is the correct result on efficient markets.
A transparent, self-pruning, MCP-native memory engine that unifies consolidation, salience-scaled adaptive forgetting, contradiction supersession, per-tenant isolation, and an exportable audit trail — evaluated on the governance properties that recall benchmarks ignore.
Things built to find something out rather than to ship.
A shadowboxing coach that names your habit a beat before you repeat it, reading punches and guard from body landmarks alone. Scored on fifteen public rounds held out by person.
A local, always-on intelligence and the machine specified to hold it. Most requests never reach a large model, but every tier has to stay resident at once, which means 384 GB of VRAM, specified from part numbers with the power ledger written down.
The multiply-accumulate at the heart of mote, carried from Verilog down to a SkyWater 130nm silicon layout: eight int8 lanes, verified against a software reference, then placed and routed to the GDSII file a foundry is handed.
How do you measure a satellite change detector's error rate without labelling anything? Point it at ground that has not moved in a century. That test found four real bugs and took false change from 7% and 39% down to 0.23% and 1.99%.
Applied LLMs
Modeling & rigor
Services & data
Ship & operate