LA LatentAtlas Email
Independent research · AI evaluation · Agent safety · Benchmark methodology

When should an AI agent act, and what do benchmark scores really show?

LatentAtlas is the independent research line of Huseyin Buldurgan. I run controlled API experiments, build Python evaluation tooling and audit benchmark source code. I publish technical reports whose results anyone can recompute from public artifacts.

  • Do tool-using language models tell relevant evidence apart from valid authorization to act?
  • Which capability claims do a benchmark’s scorers and selection rules actually support?
Open to research and research-engineering roles in AI evaluation and agent safety. Willing to relocate to the UK or EU.

Research

Six preprints, with full abstracts, directly hosted PDFs and versioned resources. These papers have not been peer-reviewed. Each one states the limits of its claims.

Browse all publications and PDFs

Benchmark audit

RE-Bench Is a Systems Benchmark: What Its Scorers and Selection Rules Actually Support

An unaffiliated audit of all seven RE-Bench task packages at a pinned source revision. It finds three measurement limitations: single-call kernel timings, scaling-law scores that are computed without running the proposed training, and a gap between the documented and implemented parameter-distance metric. It ships eight bounded contract checks. Historical results were not reproduced.

Empirical evaluation

When Relevant Evidence Is Not Permission to Act: Authority-to-Action v0.8.3

100 synthetic cases in 25 matched four-variant groups, run on two API-delivered systems for three epochs each (600 protocol-complete runs). Neither system acted in any of the 300 invalid-authority runs. Exact authorized execution differed: 150/150 versus 99/150. The results cover this synthetic protocol only; they are not production incident rates or a general model ranking.

Mathematics

Fourth-order asymptotics and monotonicity in slope-constrained moment design

A mathematical preprint on minimum-amplitude design under moment and slope constraints. It gives an explicit fourth-order value expansion for finite and countably infinite switch sets under stated hypotheses, exact examples and a regularity obstruction, and a positive theta-kernel application with uniform sixth-order error bounds and a feasible-transport proof of strict ordering.

Hüseyin Buldurgan · Version v1 · Preprint

Code & data

Public artifacts that let anyone check the reported numbers without trusting the author.

latentatlas-evidence-evals

A standard-library Python package: an Inspect evaluation, deterministic verifiers, a transcript-free 600-row results ledger, unit tests and GitHub Actions. It also hosts the RE-Bench contract checks. MIT licence.

Authority-to-Action dataset

The 100 synthetic cases and the 600-row results ledger, byte-identical to the GitHub release (same SHA-256 hashes). The dataset card carries a training-data canary string. MIT licence.

Approach

How the work is built so that its claims stay checkable.

Matched variants and negative controlsFactorial case designs isolate the variable under test. Controls show what a passing score does not mean.
Deterministic scoringTool-use outcomes are scored by code, not by model graders. Execution errors are kept separate from provider refusals.
Pinned sources and hashesAudits target fixed revisions. Released files carry SHA-256 manifests, so later readers can confirm they are checking the same artifact.
Bounded claimsEvery report says what its evidence does not establish. Unresolved results are reported as unresolved.

About

I’m Huseyin Buldurgan, an independent researcher based in Adana, Türkiye. I started LatentAtlas in January 2026 to study how AI systems should treat evidence before they act on it. The work grew out of practical data problems, where similarity-based matching kept producing confident but unsupported decisions.

Before research I spent a decade in business roles. I was a sales specialist and then sales manager at Vestel (2016–2021), onboarding investors and business partners and bringing new stores into the regional network. I also worked in investor sourcing at mbco Strategy Consulting and in a family agricultural business. I hold a BA in Management from Sabancı University.

I’m looking for a research or research-engineering role where I can work full-time on evaluation validity and agent safety, and I’m willing to relocate to the UK or EU.

Focus
AI evaluation, agent safety, benchmark methodology, evidence quality
Methods
Controlled API experiments, Inspect AI, deterministic scoring, source audits, grouped bootstrap analysis
Tools
Python 3.11, Inspect AI, provider APIs, JSON/JSONL, Git, GitHub Actions
Location
Adana, Türkiye · open to relocation

Contact

Research collaboration, role enquiries or questions about the work.