Skip to content

feat(extract): export tables and lists as JSON/CSV with provenance - #382

Open
NianJiuZst wants to merge 2 commits into
Tencent:mainfrom
NianJiuZst:codex/structured-extraction
Open

NianJiuZst wants to merge 2 commits into
Tencent:mainfrom
NianJiuZst:codex/structured-extraction

Conversation

@NianJiuZst

@NianJiuZst NianJiuZst commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Agents can already read page text and HTML, but exporting orders, reports or search
results requires site-specific scripts and a separate reconstruction step. This
adds a passive extraction primitive that returns columns, rows and their source
together. For example, an order ID such as "000123" remains a string, merged cells
retain their anchors, and a virtual grid declaring 101 rows but loading only two
returns those two rows with incomplete coverage.

Implementation

  • Add tool.extract and bsk extract discover|table|list. Discovery returns bounded
    session/document handles for semantic containers in the main page, open shadow
    roots and available frames, including OOPIFs.
  • Support HTML/ARIA tables and treegrids, hierarchical or missing headers, duplicate labels,
    explicit header IDs, spans, footer rows, row/column indices, and schema-driven
    list/card extraction with text, link or attribute fields. Reorder explicitly
    indexed ARIA rows and join unambiguous, nonoverlapping column fragments; reject
    conflicting cells.
  • Return stable column keys, original string/null values, page/frame provenance,
    row locations, spans, coverage and warnings. Bound collection by rows, columns,
    bytes, traversal work and time; only return complete rows. Cache sibling
    positions for linear locator construction, retain completed results on budget
    exhaustion, and prefer table candidates in bounded discovery.
  • Read rendered text across open shadow roots and slots, preserve block/preformatted
    boundaries, and read live input/textarea/select values. Use native cell spans,
    recognize thead cells, and report omitted hidden rows.
  • Add UTF-8 CSV export with provenance columns and a metadata sidecar. Preserve
    null positions and record the CSV SHA-256. Optional --csv-safe escapes
    formula-like values while retaining originals. Require --overwrite to replace
    existing output.
  • Expose extraction through the existing DSH browser_inspect action, using its
    session registry and normal runner/queue, including extractTimeoutMs. Add
    protocol schemas, focused unit tests and CLI/plugin references.

Document/attachment checks fence collection. Navigation, detached nodes, expired
handles and foreign sessions cannot reuse a discovered target. Cancellation
releases remote object groups. Handles retain their original tab when the active
tab changes. Extraction does not scroll, paginate or change
focus; CSV encoding and file paths stay on the CLI host.

Real-browser validation: 22 scenario groups

All 22 scenario groups passed on 2026-10-07 with the final rebuilt CLI and
extension from commit 6d413f181362b845a2b317b3f99deb16a39e7ca9, using Chrome for Testing
153.0.8010.12
on macOS / Apple Silicon. Each run used an owned browser profile,
an isolated foreground daemon and the unpacked extension.

The assertions exercised the actual CLI → daemon → extension → Chrome chain
against controlled fixtures. The runner, fixtures, results and binary hashes are
retained outside the source commits. This does not claim validation on production
customer websites.

# Scenario group Verified result
1 Container discovery Native tables, open shadow roots, same-origin frames and cross-site frames are discovered; an actual OOPIF target is asserted; hidden data is excluded.
2 Orders and provenance Three records; leading zeroes, Chinese text, commas/quotes/newlines and empty cells preserved; page/frame URL and original row positions retained; selector and target extraction agree.
3 Hierarchical headers, merged cells and footers Header paths include 营收 → 本月 / 上月; row labels stay data; merged positions are null with explicit span anchors; footer provenance is separate.
4 Nested, duplicate, headerless and empty tables Parent rows exclude nested cells; duplicate labels keep different keys; missing headers are generated; the first headerless data row is retained.
5 Partial ARIA grid Loaded source rows 51–52 returned with declared row count 101 and incomplete coverage; missing column 3 remains null.
6 Search cards and semantic lists Field schemas return titles, resolved links, summaries and null missing attributes; default semantic-list extraction also succeeds.
7 Shadow/frame target extraction Discovery handles read open-shadow, same-origin and cross-site tables with the correct page and frame provenance.
8 Row limits and invalid targets A one-row cap returns a complete record and marks truncation; ambiguous, invalid, hidden and missing targets fail explicitly.
9 CSV export and metadata CSV quoting, provenance, matching SHA-256, formula-text escaping and refusal to overwrite existing output are verified.
10 Passive behavior Page data, scroll position and focus remain unchanged after extraction.
11 Byte limits A 2048-byte budget returns three complete rows plus metadata, fits the compact JSON budget and marks truncation/incompleteness.
12 Removed nodes and navigation Handles are rejected as stale after their node is removed or the page reloads.
13 Session isolation Another session cannot reuse a discovery handle.
14 Large tables without row IDs A 6,000-row table returns 500 rows by default and 5,000 at the configured cap, with row-limit truncation.
15 Text and form values Template whitespace, block boundaries, slotted content, live controls and thead cells are handled.
16 Native spans and hidden rows Native span recovery and hidden-row coverage/warnings agree with browser behavior.
17 ARIA compatibility Treegrids, reordered row indices and unambiguous column fragments work; conflicting fragments fail.
18 Discovery limits and ordering Tables remain discoverable after navigation lists; traversal exhaustion retains collected targets with truncation.
19 Complete rows on budget exhaustion The 28-column regression returns complete accumulated rows instead of throwing away the result.
20 JSON/CSV column mapping Twelve-column output remains correctly aligned by columns[].key, independent of JSON property order.
21 Ambiguous list links Default extraction keeps item text with a null URL and warning; explicit ambiguous fields still fail.
22 Handle tab routing A saved target resolves to its original tab after another tab becomes active.

The screenshots below are retained from the original 13-scenario run. Viewport
and full-page screenshots were captured through bsk screenshot. A
separate Playwright layout check confirmed a 760 px viewport had no document
horizontal overflow and no console warnings/errors. That check is supplementary;
the extraction proof comes from the real CLI → daemon → extension chain.

Screenshots from the test run

Order and report fixtures captured through bsk screenshot:

Orders and report fixture in real Chrome

Full fixture: tables, search results, shadow DOM and frames

Full fixture captured through BrowserSkill

Supplementary layout check at 760 px (Playwright)

760 px fixture with no document horizontal overflow

Automated validation and CI

Validation rerun for the final source on 2026-10-07:

  • Chromium collector regressions: 205 passed, including 160 table shapes
    (1–40 columns, with/without headers and row IDs), a 5,000-item list, text/slot/
    control cases, invalid spans, ARIA conflicts, and traversal/byte-limit cases.
  • Additional handler regressions: 5 passed for target-tab preparation,
    expired/missing handles, frame deadlines, global table priority and partial
    discovery handles. Additional DSH timeout routing regression: 1 passed.
  • Full extension suite: 2,449 passed, 122 skipped by the existing opt-in
    configuration; includes 28 extraction normalization/lifecycle tests.
  • Full DSH plugin suite: 451 passed. CSV-focused Rust tests: 3 passed.
  • Node script suite: 13 passed, including LF and CRLF skill budgets.
  • Extension and plugin type checks/builds, CLI build, focused Biome, Rust
    formatting and Cargo/npm skill package checks passed.

The new regression runners, test cases, fixtures and result files are retained
outside the source commits. The full Rust suite was not rerun locally for this
follow-up, which changes no Rust source. All 9 hosted checks passed for
6d413f181362b845a2b317b3f99deb16a39e7ca9, including the full Rust CI job, both
Windows daemon jobs, frontend checks and code scanning. See the
PR Checks tab.

Scope and validation boundaries

Extraction reads the loaded DOM. It does not claim the entire server-side dataset
is complete, follow pagination/infinite scrolling, OCR canvas content, or discover
closed shadow roots. CSV/metadata writes are individually atomic, not one
filesystem transaction; consumers can verify the recorded CSV hash.

This validates desktop Chromium on macOS. DSH integration was tested at the
adapter/unit level; no live DSH-host session or Edge run is claimed.

See the extraction contract for the API and CSV behavior.
The code diff contains implementation, focused unit tests and usage documentation.
The HTML test pages, browser/regression runners, screenshots, logs and one-off
JSON/CSV exports are kept outside the source commits. Screenshots are attached
above; the remaining materials are preserved in the local evidence bundle.

@NianJiuZst
NianJiuZst force-pushed the codex/structured-extraction branch from 4cffd5e to b93bef1 Compare October 1, 2026 04:51

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant