Repository navigation
feat(extract): export tables and lists as JSON/CSV with provenance - #382
Open
NianJiuZst wants to merge 2 commits into
Open
NianJiuZst wants to merge 2 commits into
NianJiuZst wants to merge 2 commits into
Conversation
NianJiuZst
force-pushed
the
codex/structured-extraction
branch
from
October 1, 2026 04:51
4cffd5e to
b93bef1
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Agents can already read page text and HTML, but exporting orders, reports or search
results requires site-specific scripts and a separate reconstruction step. This
adds a passive extraction primitive that returns columns, rows and their source
together. For example, an order ID such as "000123" remains a string, merged cells
retain their anchors, and a virtual grid declaring 101 rows but loading only two
returns those two rows with incomplete coverage.
Implementation
tool.extractandbsk extract discover|table|list. Discovery returns boundedsession/document handles for semantic containers in the main page, open shadow
roots and available frames, including OOPIFs.
explicit header IDs, spans, footer rows, row/column indices, and schema-driven
list/card extraction with text, link or attribute fields. Reorder explicitly
indexed ARIA rows and join unambiguous, nonoverlapping column fragments; reject
conflicting cells.
row locations, spans, coverage and warnings. Bound collection by rows, columns,
bytes, traversal work and time; only return complete rows. Cache sibling
positions for linear locator construction, retain completed results on budget
exhaustion, and prefer table candidates in bounded discovery.
boundaries, and read live input/textarea/select values. Use native cell spans,
recognize
theadcells, and report omitted hidden rows.null positions and record the CSV SHA-256. Optional
--csv-safeescapesformula-like values while retaining originals. Require
--overwriteto replaceexisting output.
browser_inspectaction, using itssession registry and normal runner/queue, including
extractTimeoutMs. Addprotocol schemas, focused unit tests and CLI/plugin references.
Document/attachment checks fence collection. Navigation, detached nodes, expired
handles and foreign sessions cannot reuse a discovered target. Cancellation
releases remote object groups. Handles retain their original tab when the active
tab changes. Extraction does not scroll, paginate or change
focus; CSV encoding and file paths stay on the CLI host.
Real-browser validation: 22 scenario groups
All 22 scenario groups passed on 2026-10-07 with the final rebuilt CLI and
extension from commit
6d413f181362b845a2b317b3f99deb16a39e7ca9, using Chrome for Testing153.0.8010.12 on macOS / Apple Silicon. Each run used an owned browser profile,
an isolated foreground daemon and the unpacked extension.
The assertions exercised the actual CLI → daemon → extension → Chrome chain
against controlled fixtures. The runner, fixtures, results and binary hashes are
retained outside the source commits. This does not claim validation on production
customer websites.
theadcells are handled.columns[].key, independent of JSON property order.The screenshots below are retained from the original 13-scenario run. Viewport
and full-page screenshots were captured through
bsk screenshot. Aseparate Playwright layout check confirmed a 760 px viewport had no document
horizontal overflow and no console warnings/errors. That check is supplementary;
the extraction proof comes from the real CLI → daemon → extension chain.
Screenshots from the test run
Order and report fixtures captured through
bsk screenshot:Full fixture: tables, search results, shadow DOM and frames
Supplementary layout check at 760 px (Playwright)
Automated validation and CI
Validation rerun for the final source on 2026-10-07:
(1–40 columns, with/without headers and row IDs), a 5,000-item list, text/slot/
control cases, invalid spans, ARIA conflicts, and traversal/byte-limit cases.
expired/missing handles, frame deadlines, global table priority and partial
discovery handles. Additional DSH timeout routing regression: 1 passed.
configuration; includes 28 extraction normalization/lifecycle tests.
formatting and Cargo/npm skill package checks passed.
The new regression runners, test cases, fixtures and result files are retained
outside the source commits. The full Rust suite was not rerun locally for this
follow-up, which changes no Rust source. All 9 hosted checks passed for
6d413f181362b845a2b317b3f99deb16a39e7ca9, including the full Rust CI job, bothWindows daemon jobs, frontend checks and code scanning. See the
PR Checks tab.
Scope and validation boundaries
Extraction reads the loaded DOM. It does not claim the entire server-side dataset
is complete, follow pagination/infinite scrolling, OCR canvas content, or discover
closed shadow roots. CSV/metadata writes are individually atomic, not one
filesystem transaction; consumers can verify the recorded CSV hash.
This validates desktop Chromium on macOS. DSH integration was tested at the
adapter/unit level; no live DSH-host session or Edge run is claimed.
See the extraction contract for the API and CSV behavior.
The code diff contains implementation, focused unit tests and usage documentation.
The HTML test pages, browser/regression runners, screenshots, logs and one-off
JSON/CSV exports are kept outside the source commits. Screenshots are attached
above; the remaining materials are preserved in the local evidence bundle.