i do engineering

PROJECTS

sotto

CASE STUDY →Sep 2026

Slower, calmer synthetic speech without slowing the words: render each sentence once at natural speed, then lengthen only the silences the model already made. Every register hits its words-per-second target at speed 1.00.

Python
FastAPI
Kokoro TTS
Self-hosted

mote

CASE STUDY →Aug 2026

A transformer inference engine in plain C with no libraries underneath, small enough to run a real chat model offline on a phone in under 90 KB of WebAssembly. Along the way: why a smaller quantization error can make the model worse.

Embedded C
Transformer
Quantization
Open Source

ràdá

CASE STUDY →Aug 2026

A pocket 60.5 GHz radar that measures movement in micrometres by reading the phase of the returning wave, enough to count a heartbeat through still air or recover sound from a surface it never touches.

60.5 GHz radar
nRF52840
Embedded C
KiCad

àyò

CASE STUDY →Jul 2026

A research platform for systematic trading in which Claude designs strategies but never writes code. Each strategy is a frozen, hashed artifact run by a deterministic engine, so the results are reproducible even though the author is not, and every idea has to clear an out-of-sample gate that tightens as more are tried.

Python
Claude
Alpaca
pandas

dimeji

CASE STUDY →Jul 2026

A typed graph for modelling an emotional history honestly. Every claim has to declare how it is known, hypotheses are never promoted to facts, and contradictions stay open instead of being resolved.

Framework
React Flow
TypeScript
Epistemics

vibe

CASE STUDY →Jul 2026

A Claude Code skill that holds the assistant in one persona for a whole session while the code, comments and commits stay clean. One Markdown file, open source.

Claude Code
Skill
Open Source
Markdown

the scribe

LIVE SITE ↗Jun 2026

A ghostwriting studio for ministry authors where the AI editor cannot invent: every margin note must quote the manuscript verbatim, and a failed check goes back to the model so it corrects itself in the same loop.

TypeScript
Next.js
OpenAI
Postgres
Prisma

RESEARCH & PAPERS

Most of my research asks one question: how do you know an automated system is right? My answer is to measure it where the truth is already known: satellite sites certified not to have changed, proof claims that either were or were not checked, a trading gate that has to reject strategies which only looked good in-sample. The first two are the main papers. The rest are shorter: a technical note and system reports on the tools behind that work, a preprint on how many projects to start, and one essay.

Faster Than They Can Be Checked

Terence Tao predicted that AI-generated proofs would accumulate faster than they could be verified. This measures it. Every public proof claim on erdosproblems.com from July to September 2026, 251 of them and 97% made with AI, was traced to see whether anyone ever checked it. Of the 224 on problems still open, 48% drew any comment, 19% reached any verdict and 4% were accepted by the site's moderators. Verdicts arrive within a day or not at all: after 60 days, 82% of claims still have none. Attaching a Lean proof made no difference, 19 people delivered every verdict, and one complete, formally verified proof was dismissed with a mistaken "already solved" that nobody corrected.

PUBLISHED: Sep 25, 2026Preprint

What Does a Change Detector Find Where Nothing Changed?

You cannot label the earth, so nobody measures how often a satellite change detector invents change. This runs the pipeline over ground another community already certified as invariant, where every detection is a false alarm by construction. The field has not done this because it cannot: on an empty ground truth, F1 and IoU are undefined for a perfect result and identically zero for every imperfect one, so they cannot tell one false positive from five hundred thousand. On the six CEOS-endorsed desert sites, a threshold chosen from the data reports a median 21.41% of certified-stable ground as changed, against 0.04% with an absolute floor. Otsu calls 39% of an empty scene changed, eases only to 27% as a real signal strengthens, then falls 26.97 points across a single 0.25-sigma step, so behaviour on scenes containing change predicts nothing about scenes without it. Seven defects surfaced, five invisible in the outputs, including a pipeline that reports thirty times less change once a large real structure appears.

PUBLISHED: Sep 23, 2026Preprint

We Are All Doppelgängers of Ourselves

The doppelgänger of folklore terrifies because it is a copy of you that is not you. But the body is a pattern rebuilt from new matter, memory is re-saved each time it is opened, and everyone who knows you carries a version of you you will never meet. From Goethe's consoling double and the brain's misfiled selves to Yoruba twin figures and a chatbot built from a dead man's messages, an essay on why there was never an original, and what the copies owe each other.

PUBLISHED: Sep 25, 2026Essay

How Much of a Small Chip Actually Computes?

A synthesized netlist is pure logic, but a layout a foundry could build is logic plus the machinery that lets it survive fabrication: antenna diodes, well taps, timing buffers, fill. Taking one signed int8 multiply-accumulate through the open Sky130 flow at lane counts from 1 to 64, I decompose every placed layout and watch the arithmetic share fall from 59 percent of the cells to 48, crossing below half, because the overhead outgrows the logic and its pieces scale against wiring and die area rather than the gates.

PUBLISHED: Aug 29, 2026Technical note

How Many Things Can You Finish? Two Effort Thresholds and Optimal Attempt Counts

An attempt needs two separate efforts: one to make it work and one to put it in front of anyone. The optimal number of simultaneous attempts depends on their sum, but the heuristic builders actually use depends on build cost alone. Agent assistance collapsed build cost and left shipping cost untouched, so the two prescriptions have diverged. Simulation plus a 48-repository archive where every stalled project is stalled at a step building cannot resolve.

PUBLISHED: Aug 21, 2026Preprint

How Many Strategies Did You Really Try? Effective Trial Counts for LLM-Proposed Backtests

When a language model proposes trading strategies, the trial count that selection-bias corrections depend on is not the number of proposals but the effective number of independent ones. Borrowing the effective-number-of-tests estimator from statistical genetics gives a principled M_eff, and on the live ÀYÒ pipeline the size of the correction turns out to be a property of the prompting protocol: diverse angles yield near-independent proposals that need no correction, while drawing repeatedly from a single narrow angle collapses fourteen proposals to three effective trials and inflates the acceptance bar enough to reject genuine strategies.

PUBLISHED: Jul 29, 2026System report

Trials-Adjusted Out-of-Sample Gating for LLM-Proposed Trading Strategies

A validation gate built to be un-foolable by construction: strategies arrive as frozen, content-hashed data artifacts, as an LLM would propose them, and a deterministic verifier trades only what survives walk-forward out-of-sample testing, a trials-adjusted acceptance bar, and robustness sweeps. Tested here on scripted strategies, on real markets it turned an in-sample Sharpe of 3.07 into a −4.90 rejection and unmasked a +0.13 out-of-sample "edge" as a 0-of-12 fluke — deploying nothing, which is the correct result on efficient markets.

PUBLISHED: Jul 20, 2026System report

Cortex: A Governed Long-Term Memory Engine for LLM Agents

A transparent, self-pruning, MCP-native memory engine that unifies consolidation, salience-scaled adaptive forgetting, contradiction supersession, per-tenant isolation, and an exportable audit trail — evaluated on the governance properties that recall benchmarks ignore.

PUBLISHED: Jul 19, 2026System report

EXPERIMENTS

Things built to find something out rather than to ship.

ika

CASE STUDY →Sep 2026

A shadowboxing coach that names your habit a beat before you repeat it, reading punches and guard from body landmarks alone. Scored on fifteen public rounds held out by person.

MediaPipe
PyTorch
Python
On-device

node

CASE STUDY →Aug 2026

A local, always-on intelligence and the machine specified to hold it. Most requests never reach a large model, but every tier has to stay resident at once, which means 384 GB of VRAM, specified from part numbers with the power ledger written down.

Threadripper PRO
384 GB VRAM
Local inference
Hardware

mak

CASE STUDY →Aug 2026

The multiply-accumulate at the heart of mote, carried from Verilog down to a SkyWater 130nm silicon layout: eight int8 lanes, verified against a software reference, then placed and routed to the GDSII file a foundry is handed.

Verilog · RTL
Sky130
OpenLane
Silicon

jùwọ̀n

CASE STUDY →Aug 2026

How do you measure a satellite change detector's error rate without labelling anything? Point it at ground that has not moved in a century. That test found four real bugs and took false change from 7% and 39% down to 0.23% and 1.99%.

Sentinel-2 · SAR
NumPy
Python
Calibration

TECHNICAL SKILLS

LANGUAGES

  • TypeScript / JavaScript (React, Next.js, Node)
  • Python (AI/ML, FastAPI & Django backends, data)
  • Swift (SwiftUI, shipped iOS apps)

AI / LLM ENGINEERING

Applied LLMs

  • Claude & OpenAI in production — structured outputs, streaming, tool-use
  • RAG, agents & multi-step tool-calling, MCP
  • Voice AI (ElevenLabs conversational agents)

Modeling & rigor

  • LLM pretraining & fine-tuning (from scratch)
  • Eval harnesses, guardrails & deterministic verification
  • PyTorch, Transformers, classic ML

FRONTEND

  • React & Next.js (App Router, RSC, streaming SSR)
  • SwiftUI (native iOS, widgets, Live Activities)
  • Tailwind CSS, design systems, motion

BACKEND & INFRA

Services & data

  • FastAPI, Django, Node — REST + SSE, background jobs (Redis/Celery)
  • PostgreSQL, Prisma, SQLite, pgvector
  • Stripe / RevenueCat billing, auth (JWT, OAuth)

Ship & operate

  • Docker, Vercel, Railway, Render, Cloudflare
  • CI, eval-gated deploys, observability