AI Systems, Model Evaluation & AI Assurance: independent research, empirical auditing, and operational tools.
EU - UK · nmairesearch.github.io · ORCID 0009-0003-4213-7769 · Zenodo Archive · Hugging Face · Email
Research focuses on designing, evaluating, and auditing AI systems from first principles: finding where AI workflows break, testing model confidence against ground truth, and building reproducible data pipelines and zero-dependency analytical tools.
Projects publish methods, data and verification scripts with project-specific coverage. Reproducing stored calculations does not establish that every source supports its claim. Historical results, corrections and unresolved checks are identified in the linked records.
Earlier background is ten years in operations and process design, quality assurance, and KPI-driven performance evaluation. BA in Business Management (2024). Bilingual English and Hungarian (C2).
Sovereign Watch. A scheduled regulatory intake supporting the AI News Board, developed with AI-assisted implementation. Source failures, pending assessments and recovery states remain visible. The operations case distinguishes project ownership from the assistant's engineering work, and tested controls from unmeasured accuracy or time savings. Case study · Public board
-
AI Infrastructure: Commitments, Delivery and Access. Examines selected dated commitments, delivery stages and access obligations. The v7.1 bundle recalculates named balances and regenerates selected numerical content and tables from frozen inputs. An evidence lookup carries source references and qualifications alongside each quantity; arithmetic and file-integrity checks do not establish source truth or causal validity.
Paper · Runnable bundle and example -
What Actually Admits a Document to FineWeb-Edu. Audited published quality-filter scores in a sample of 84,005,795 documents from 94 Common Crawl snapshots. Under the recorded rounding rule, admission requires a raw score strictly above 2.5; the sample contains 458,461 documents exactly at the excluded tie point. The classifier has a 510-token content budget. The corrected source and explorer distinguish score-distribution results from tokenisation estimates and leave the historical fixed-score extension unavailable pending recalculation. The linked historical PDF and dataset have not yet received this correction.
Interactive Tool · Zenodo DOI 10.5281/zenodo.21740081 · Dataset on Hugging Face · Source Code -
The Model Is a Dependency: Provenance and Error Detection in Financial Data Extraction. The corrected v1.1 release distinguishes model-run provenance from answer checking in financial data extraction. Selected calculations and tables are reconstructed from stored outputs, with a corrected source label and explicit historical input limits. It reports open-weight experiments only; no new model runs or deployment validation are claimed.
Interactive Tool · Zenodo DOI 10.5281/zenodo.21543579 · Source Code -
Public evidence under Article 50: a historical snapshot of twelve consumer AI products. The corrected v1.1 paper preserves a historical public-evidence snapshot of twelve consumer AI products. It identifies unresolved cells, capture-state limits and the difference between a bounded negative search and an established absence. The linked explorer retains the historical scores. This is not a current compliance assessment or a completed repeat survey.
Interactive Explorer · Zenodo DOI 10.5281/zenodo.21819102 · Source Code -
AI Infrastructure: Power, Water, and Grid Queues. Compared measured and planned water figures with explicit on-site and electricity-related boundaries. The power study distinguishes historical PJM generation-queue completion from data-centre load requests. Its generation result is not a measured completion rate for AI data centres or the UK queue. Water relocation estimates depend on the stated cooling and electricity assumptions.
Water Tracker · Zenodo DOI 10.5281/zenodo.21318960 · Power Demand Tool -
The CEO Pay-vs-Delivery Scorecard. Built a dated issuer-level view of SEC Pay Versus Performance disclosures, separating Summary Compensation Table totals from Compensation Actually Paid. CAP is an accounting fair-value remeasurement, not cash received. Board-target cases and peer-relative returns are separate descriptive comparisons, without attributing returns to the CEO.
Scorecard · Zenodo DOI 10.5281/zenodo.20680108 · Source Code
- AI Systems & Evaluation: Agentic workflows, tool-calling interfaces and MCP connectors, failure-mode discovery, uncertainty calibration (token entropy), prompt boundary testing, and human-in-the-loop verification.
- Model Governance & Assurance: EU AI Act public-evidence assessment (Article 50), Model Risk Management (MRM), content provenance (C2PA, machine-readable markings), and ISO/IEC 42001 concepts.
- Quantitative Research & Data: Dataset curation audits (FineWeb-Edu), primary-record pipelines (SEC EDGAR and XBRL parsing), reproducible build.py and reproduce.py workflows, and compute and grid infrastructure modelling.
- Operations & Quality Assurance: Ten years in operational process architecture, operational health-and-safety audit preparation (reported 92 per cent external audit score for the night operation), root-cause analysis, and cross-functional standard operating procedures.
- Interactive portfolio: https://nmairesearch.github.io/
- Curriculum vitae: awaiting upload
- Zenodo archive: https://zenodo.org/search?q=metadata.creators.person_or_org.name%3A%22NM%20AI%20Research%22
- Hugging Face datasets: https://huggingface.co/NMAIResearch
- ORCID record: https://orcid.org/0009-0003-4213-7769
- Contact: NMAIResearch@proton.me
Verification: the summaries preserve each project's stated measurement and reproduction scope. The linked records distinguish historical results, corrected outputs and unresolved work. A passing calculation check does not establish semantic support for every claim.
AI disclosure: the author directs the research questions, methods and substantive decisions. AI models assist research, drafting and implementation. OpenAI GPT-6 assisted these summary corrections; OpenAI and other assisting providers are subjects or competitors in parts of the portfolio. Project-specific disclosures identify their contributions. Independent review and author acceptance of this correction remain separate.