Mina Maged Zekry Gayid

Research and evidence gates

Reliable agentic AI for high-stakes systems.

Reliable Agentic AI for High-Stakes Systems

My research direction connects agent failure analysis, source-aware memory and evidence-bound review. The immediate work is local evaluation tooling and controlled simulations. Meaningful deployment needs independently labeled tasks, trustworthy traces and authenticated human decisions.

ELLM: a rejected adapter is useful evidence

The existing ELLM experiment uses a pinned SmolLM2-135M-Instruct base and a LoRA adapter for synthetic four-label policy triage. The held-out result was 79/256 exact answers (30.9%). It released 177/192 non-release cases (92.2%) and failed its safety promotion gate. This adapter must not be used as an operational approval authority.

The deterministic evidence gate is a separate component; its behavior cannot be credited to the model. Training manifests, split hashes, runtime tests and failed results are available in the ELLM model repository. The base is an existing English model; no Arabic benchmark or training from scratch is claimed.

Clinical and scientific evidence gates

Dental quality ratings currently include model-generated weak labels. Expert references, real patient-disjoint splits and external validation are needed. Neuroimaging overlap and calibration functions need independent licensed labels before reporting performance. Source hashes in scientific retrieval identify content provenance, not scientific correctness.

Simulation before hardware

The robotics fault report exercises rejection and stop decisions in a simulator. It does not establish physical collision safety or medical-device readiness. The next work requires richer perturbation sweeps, timing analysis and independent review.

Read the engineering note → · Implementation ledger