Skip to content
View NMAIResearch's full-sized avatar

Block or report NMAIResearch

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
NMAIResearch/README.md

NM AI Research

AI Systems, Model Evaluation & AI Assurance: independent research, empirical auditing, and operational tools.

EU - UK · nmairesearch.github.io · ORCID 0009-0003-4213-7769 · Zenodo Archive · Hugging Face · Email


Research focuses on designing, evaluating, and auditing AI systems from first principles: finding where AI workflows break, testing model confidence against ground truth, and building reproducible data pipelines and zero-dependency analytical tools.

Projects publish methods, data and verification scripts with project-specific coverage. Reproducing stored calculations does not establish that every source supports its claim. Historical results, corrections and unresolved checks are identified in the linked records.

Earlier background is ten years in operations and process design, quality assurance, and KPI-driven performance evaluation. BA in Business Management (2024). Bilingual English and Hungarian (C2).


Operations evidence

Sovereign Watch. A scheduled regulatory intake supporting the AI News Board, developed with AI-assisted implementation. Source failures, pending assessments and recovery states remain visible. The operations case distinguishes project ownership from the assistant's engineering work, and tested controls from unmeasured accuracy or time savings. Case study · Public board

Selected Work & Empirical Audits

  • AI Infrastructure: Commitments, Delivery and Access. Examines selected dated commitments, delivery stages and access obligations. The v7.1 bundle recalculates named balances and regenerates selected numerical content and tables from frozen inputs. An evidence lookup carries source references and qualifications alongside each quantity; arithmetic and file-integrity checks do not establish source truth or causal validity.
    Paper · Runnable bundle and example

  • What Actually Admits a Document to FineWeb-Edu. Audited published quality-filter scores in a sample of 84,005,795 documents from 94 Common Crawl snapshots. Under the recorded rounding rule, admission requires a raw score strictly above 2.5; the sample contains 458,461 documents exactly at the excluded tie point. The classifier has a 510-token content budget. The corrected source and explorer distinguish score-distribution results from tokenisation estimates and leave the historical fixed-score extension unavailable pending recalculation. The linked historical PDF and dataset have not yet received this correction.
    Interactive Tool · Zenodo DOI 10.5281/zenodo.21740081 · Dataset on Hugging Face · Source Code

  • The Model Is a Dependency: Provenance and Error Detection in Financial Data Extraction. The corrected v1.1 release distinguishes model-run provenance from answer checking in financial data extraction. Selected calculations and tables are reconstructed from stored outputs, with a corrected source label and explicit historical input limits. It reports open-weight experiments only; no new model runs or deployment validation are claimed.
    Interactive Tool · Zenodo DOI 10.5281/zenodo.21543579 · Source Code

  • Public evidence under Article 50: a historical snapshot of twelve consumer AI products. The corrected v1.1 paper preserves a historical public-evidence snapshot of twelve consumer AI products. It identifies unresolved cells, capture-state limits and the difference between a bounded negative search and an established absence. The linked explorer retains the historical scores. This is not a current compliance assessment or a completed repeat survey.
    Interactive Explorer · Zenodo DOI 10.5281/zenodo.21819102 · Source Code

  • AI Infrastructure: Power, Water, and Grid Queues. Compared measured and planned water figures with explicit on-site and electricity-related boundaries. The power study distinguishes historical PJM generation-queue completion from data-centre load requests. Its generation result is not a measured completion rate for AI data centres or the UK queue. Water relocation estimates depend on the stated cooling and electricity assumptions.
    Water Tracker · Zenodo DOI 10.5281/zenodo.21318960 · Power Demand Tool

  • The CEO Pay-vs-Delivery Scorecard. Built a dated issuer-level view of SEC Pay Versus Performance disclosures, separating Summary Compensation Table totals from Compensation Actually Paid. CAP is an accounting fair-value remeasurement, not cash received. Board-target cases and peer-relative returns are separate descriptive comparisons, without attributing returns to the CEO.
    Scorecard · Zenodo DOI 10.5281/zenodo.20680108 · Source Code


Core Focus & Methods

  • AI Systems & Evaluation: Agentic workflows, tool-calling interfaces and MCP connectors, failure-mode discovery, uncertainty calibration (token entropy), prompt boundary testing, and human-in-the-loop verification.
  • Model Governance & Assurance: EU AI Act public-evidence assessment (Article 50), Model Risk Management (MRM), content provenance (C2PA, machine-readable markings), and ISO/IEC 42001 concepts.
  • Quantitative Research & Data: Dataset curation audits (FineWeb-Edu), primary-record pipelines (SEC EDGAR and XBRL parsing), reproducible build.py and reproduce.py workflows, and compute and grid infrastructure modelling.
  • Operations & Quality Assurance: Ten years in operational process architecture, operational health-and-safety audit preparation (reported 92 per cent external audit score for the night operation), root-cause analysis, and cross-functional standard operating procedures.

Links & Verification


Verification: the summaries preserve each project's stated measurement and reproduction scope. The linked records distinguish historical results, corrected outputs and unresolved work. A passing calculation check does not establish semantic support for every claim.

AI disclosure: the author directs the research questions, methods and substantive decisions. AI models assist research, drafting and implementation. OpenAI GPT-6 assisted these summary corrections; OpenAI and other assisting providers are subjects or competitors in parts of the portfolio. Project-specific disclosures identify their contributions. Independent review and author acceptance of this correction remain separate.

Popular repositories Loading

  1. ceo-pay-scorecard ceo-pay-scorecard Public

    Interactive front-end to the CEO Pay-vs-Delivery Scorecard: granted pay against delivered performance for 495 of the 503 S&P 500 issuers, built from SEC Pay-versus-Performance disclosures on EDGAR.…

    HTML

  2. forecast-scorecard forecast-scorecard Public

    Interactive front-end to the AI Energy-Demand Forecast Scorecard: how published forecasts of data-centre electricity demand disperse, how they are revised, and whether they can be reproduced from w…

    HTML

  3. announced-vs-deliverable-power announced-vs-deliverable-power Public

    Interactive front-end and reproducible models for Announced vs. Deliverable AI Power Demand. DOI 10.5281/zenodo.20559430

    HTML

  4. NMAIResearch.github.io NMAIResearch.github.io Public

    Landing page for NM AI Research. Interactive tools and reproducible studies on the financing, energy use and market conduct of the AI build-out. Each study publishes its data, a script that regener…

    HTML

  5. ai-water-tracker ai-water-tracker Public

    Rebuilds the Ren et al. Water Consumption Impact index for AI data centres from public data, and adds the hydropower-coupling and off-site relocation channels the published index leaves open. Water…

    HTML

  6. bot-energy-bound bot-energy-bound Public

    An order-of-magnitude bound on the electricity used by automated web traffic, and an account of why public data supports no single figure. Two methods, both reported, disagreeing by a factor of 5.3…

    Python