Research

I conduct independent research on AI agents, evaluation, and the tools models use to inspect evidence. Unfinished work is labeled by its current status.

Working paper

Evidence preservation in model-facing tools

Tests whether document tools preserve the evidence a model needs to support an answer. The experiment compares source documents with the text and metadata exposed to an agent.

Archive tools · evaluation · evidence

Completed pilot

Social selection among language-model agents

The pilot tested whether agents learn to cooperate when partners can refuse to work with them. One bounded training run reduced cooperation. The experiment did not show reliable learning or transfer, and the negative result now defines the next design question.

Multi-agent learning, cooperation, negative result

Deployed system

Research system for the Piłsudski archive

A multilingual research agent for the Józef Piłsudski Institute of America. It searches 13,000 documents and 87,000 page scans through full-text search, text embeddings, image embeddings, and eleven agent tools.

Historical archives · retrieval · research infrastructure