Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
64 changes: 61 additions & 3 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,7 @@ lighton/
workspace.py # Workspace, active-record, lives at root
apikey.py # ApiKey / ApiKeyScope, active-record, lives at root
tag.py # Tag, active-record (list/create/delete only; no single GET)
content_type.py # ContentType/Facet/Attribute, content-type taxonomy + file facets
content_type.py # ContentType/Facet/Attribute/FacetAction/FacetResult, taxonomy + file facets
file.py # File, active-record + wait_all(); upload = ingestion
batch.py # ingest_many() batch upload behavior: BatchIngestJob (threads/poll)
job.py # ParseJob/ExtractJob, client-bound async handles you poll()
Expand Down Expand Up @@ -152,6 +152,16 @@ parsed), `ServerError` (5xx), and `MaintenanceError`.
payload stays on `.body`. A 503 **without** the marker stays a plain `ServerError`
(pinned by a test). Still not retried: 5xx never is, and a maintenance window outlasts
any cooldown worth sleeping through.
- **`LightOnAPIError.index`** is the 0-based position of the failing action inside a
batch request, `None` otherwise. Parsed in the **base `__init__`** by `_index()`
(sibling of `_retry_after`/`_timestamp`), so every subclass inherits it through its
`super().__init__` and `from_response` stays the single construction point; the
position is also appended to the message (`... (action 3)`) so a bare traceback
names the offender. Not a `BatchActionError` subclass, because the API sends
`index` on **400/403/404/422 alike**: a body-keyed class would either break
`except NotFoundError` for batch callers or fork four ways for one integer, and
`MaintenanceError` stays the *one* body-keyed mapping. `_index()` tests
`isinstance(value, int)` rather than truthiness, since index 0 is a real answer.

## Resource management: active-record

Expand Down Expand Up @@ -349,8 +359,56 @@ use `_ActiveRecord.list`.
`File` classification (all via `POST /files/<id>/facets` with an `action`): `classify`/
`unclassify` (assign/remove a content type, T2), `set_attribute`/`clear_attribute` (an
attribute value under an assigned type, T3), each accepting a `ContentType` or a path
string (one `_facet(action, ct, **extra)` helper builds the body). `facets()` GETs the
file's assigned types as `list[Facet]`. Like tags, File models no facet fields locally.
string. `facets()` GETs the file's assigned types as `list[Facet]`. Like tags, File
models no facet fields locally.

- **`batch_facets(actions)`** is the batch sibling (POST `/files/<id>/facets/batch`,
max **50**), and unlike `ContentType.batch` it is **typed both ways**. That
divergence is the point, not an oversight: a taxonomy action spans five different
shapes (`adopt` takes a path list, `define_content_type` code/label/parent,
`define_attribute` seven fields), so modelling it means five models or one wide
model whose valid fields depend on the action; a file-facet action has exactly
**one** shape (four fields, four verbs), which `FacetAction` models cleanly. Raw
dicts are still accepted alongside `FacetAction`, so the escape hatch is what the
two surfaces share. `ContentType.batch` was deliberately left untouched: it is
shipped public API and retyping its return would break `r["status"]` for everyone.
`FacetAction` is the one curated model with **`extra="forbid"`**: `ignore` suits
read models (drop response noise), but on a write model it would silently drop a
misspelled field from the body, so a typo raises instead (a test pins it). Unmodelled
fields go through a raw dict.
- **Naming split.** The `FacetAction` constructors carry the **SDK's** method names
(`classify`/`unclassify`/`set_attribute`/`clear_attribute`) so a batch is a
mechanical transcription of the one-by-one calls it replaces; `FacetActionType`
(enums.py, its full domain is documented, mirrors the generated
`FileFacetActionRequestActionEnum`) and the wire carry the **API's**
(`set_value`/`clear_value`). The enum exists partly to make that mapping
discoverable. Method name is `batch_facets`, not `facets_batch`: every write on
`File` is verb-first, and bare `batch` is unusable because `Workspace`/`batch.py`
already own a batch *ingest* concept.
- **One wire encoder.** `FacetAction._body()` is the single place that knows the
body field names, and `_facet` now takes a `FacetAction`, so the four single-action
methods and the batch cannot drift; a test pins single-action bodies equal to batch
bodies. `FacetAction`/`FacetResult`/`MAX_FACET_ACTIONS` live in `content_type.py`
next to `Facet`/`Attribute` (the `Template` precedent for taxonomy-adjacent data
models); `types/` would invert the dependency by pulling `ContentType` down into it.
- **The 50 cap raises a `ValueError` client-side**, the same trade as
`define_attribute`'s missing-`choices` check, and empty is a local no-op (the
`tag`/`untag`/`delete_many` convention). It is deliberately **not chunked**: the
endpoint fails fast at the offending action and commits everything before it, so
chunking would both scatter that boundary (the reported `index` would be relative
to a batch the caller never wrote) and forfeit the single BM25 reindex that is the
whole reason the endpoint exists.
- **`FacetResult` is curated, not the generated `BatchResultItem`**: the generated
name is anonymous and can't document which status belongs to which verb. Its `data`
stays a raw dict for the reason `ContentType.batch` already gives (a classification
for `classify`, an attribute for `set_value`, null for the 204 verbs, so no one
model fits), but the *envelope* is uniform, which is what `FacetResult` buys. No
back-reference to the action: results only come back on success, complete and in
order, so `results[i]` is `actions[i]` by construction.
- **Partial commit, unlike `delete_many`.** `/files/bulk-delete` is all-or-nothing
server-side, so it raises and there is no per-item report to give. This endpoint
commits the prefix before the failure and does **not** return those results, so the
*position* is the payload, and it rides on the exception (`LightOnAPIError.index`).

If adding new resources, subclass `_ActiveRecord`: set `_base`/`_resource`, declare the
field schema (narrow `id`), and add `create()`/`save()`. Everything else is inherited.
Expand Down
70 changes: 70 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -849,6 +849,70 @@ doc.clear_attribute("legal:contract:nda", "jurisdiction")
doc.unclassify("legal:contract:nda")
```

Doing several of these at once? See
[Many writes in one request](#many-writes-in-one-request).

### Many writes in one request

Classifying a document is rarely one call: it's the `classify`, then one write per
attribute. `batch_facets()` sends up to 50 of them in a **single request**, which
also reindexes the document once instead of once per action, and that is the part
that actually costs time. Build the list with `FacetAction`, whose constructors
take the same arguments as the methods above, so a batch is a transcription of the
calls it replaces:

```python
from lighton import FacetAction, File

doc = File.get_by_name(client, "nda-2026.pdf", workspace=42)[0]

results = doc.batch_facets([
FacetAction.classify("legal:contract:nda"),
FacetAction.set_attribute("legal:contract:nda", "jurisdiction", "FR"),
FacetAction.set_attribute("legal:contract:nda", "signed_on", "2026-07-01"),
FacetAction.set_attribute("legal:contract:nda", "signed", True),
])
print([r.status for r in results]) # [201, 201, 201, 201]
```

The list is inert until it reaches a file, so the same one applies to a whole
corpus:

```python
actions = [
FacetAction.classify("legal:contract:nda"),
FacetAction.set_attribute("legal:contract:nda", "jurisdiction", "FR"),
]
for doc in File.list(client, workspace_id=42, extension="pdf"):
doc.batch_facets(actions)
```

You get one `FacetResult` per action, in the order sent, so `results[i]` belongs to
`actions[i]`. Each carries a `status` (201 created, 200 already applied, 204 for
`unclassify`/`clear_attribute`) and `data`, which is `None` for those 204 actions.

**When one action fails.** The batch is not transactional. A malformed action is
caught up front and nothing runs, but a *domain* error (an unknown content type,
setting a value before classifying, a sibling conflict) stops at that action and
leaves everything before it applied. The error carries the 0-based position, so
the offender is addressable, and every action is idempotent, so the fix is to
correct it and resend the whole list:

```python
from lighton.exceptions import LightOnAPIError

try:
doc.batch_facets(actions)
except LightOnAPIError as e:
print("failed at action", e.index, e.body["detail"])
print("already applied:", actions[:e.index])
```

`index` is `None` on any non-batch error. Past 50 actions `batch_facets()` raises a
`ValueError` before the request goes out: split the list yourself rather than have
the SDK guess where to cut, since chunking would restore the per-chunk reindex the
endpoint exists to avoid.

### Building the taxonomy

Starting from nothing? Adopt a starter tree from the catalog:
Expand Down Expand Up @@ -898,6 +962,12 @@ results = ContentType.batch(client, [
print([r["status"] for r in results]) # [201, 201]
```

This is the taxonomy-side sibling of
[`batch_facets()`](#many-writes-in-one-request). The entries stay raw dicts here
because the taxonomy spans five different action shapes (`adopt` takes a path list, `define_content_type`
takes code/label/parent, `define_attribute` takes seven fields), whereas a file
facet action has exactly one shape, which is what `FacetAction` models.

## API keys

Same active-record style. The plaintext secret is available **only** right after `create()`.
Expand Down
15 changes: 14 additions & 1 deletion lighton/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -3,11 +3,20 @@
from lighton._client import LightOn
from lighton.apikey import ApiKey, ApiKeyScope
from lighton.batch import BatchIngest, BatchIngestJob, BatchProgress, FailedIngest
from lighton.content_type import Attribute, ContentType, Facet, Template
from lighton.content_type import (
MAX_FACET_ACTIONS,
Attribute,
ContentType,
Facet,
FacetAction,
FacetResult,
Template,
)
from lighton.enums import (
AttributeType,
DownloadPurpose,
ExecMode,
FacetActionType,
FileStatus,
JobStatus,
RelevanceScoring,
Expand Down Expand Up @@ -54,12 +63,16 @@
"ExternalMetadata",
"ExtractJob",
"Facet",
"FacetAction",
"FacetActionType",
"FacetResult",
"FailedIngest",
"File",
"FileStatus",
"JobStatus",
"LightOn",
"LightOnConfiguration",
"MAX_FACET_ACTIONS",
"ParseJob",
"RelevanceScoring",
"ReprocessLevel",
Expand Down
164 changes: 161 additions & 3 deletions lighton/content_type.py
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,8 @@

`Facet` is a content type *assigned to a file* together with the file's attribute
values on it (see `File.classify()` / `File.facets()`). `Attribute` is the shared
name/type/value shape used by both.
name/type/value shape used by both. `FacetAction`/`FacetResult` are the write and
result shapes of `File.batch_facets()`.
"""

from __future__ import annotations
Expand All @@ -16,16 +17,19 @@
from builtins import list as _list
from typing import TYPE_CHECKING, Any

from pydantic import BaseModel, ConfigDict, Field
from pydantic import BaseModel, ConfigDict, Field, model_validator

from lighton.enums import AttributeType
from lighton.enums import AttributeType, FacetActionType
from lighton.utils import _compact, _path

if TYPE_CHECKING:
from lighton._client import LightOn

_BASE = "/api/v3/content-types"

MAX_FACET_ACTIONS = 50
"""Actions per `File.batch_facets()` call, the API's cap. Split a longer job yourself."""


class Attribute(BaseModel):
"""One attribute of a content type, a definition, or a value set on a file.
Expand Down Expand Up @@ -320,5 +324,159 @@ class Facet(BaseModel):
)


class FacetAction(BaseModel):
"""One classification write, applied by `File.batch_facets()`.

Build these with the constructors, not the fields: each takes the same
arguments in the same order as the `File` method of the same name, so a batch
is a transcription of the single-action calls it replaces.

doc.batch_facets([
FacetAction.classify(nda),
FacetAction.set_attribute(nda, "jurisdiction", "FR"),
])

Nothing is sent until the list reaches a file, so one list applies to many.
"""

# forbid, not ignore: this is a write model, a misspelled field must fail
# loudly rather than vanish from the body. Raw dicts are the unmodelled escape.
model_config = ConfigDict(extra="forbid")

action: FacetActionType = Field(
description="The write to perform, the API's verb (see FacetActionType)."
)
content_type_path: str = Field(
description="Content type the action applies to, e.g. legal:contract:nda."
)
attribute_name: str | None = Field(
None,
description="Attribute identifier in snake_case; required by the value verbs.",
)
value: Any = Field(
None,
description=(
"Value for set_attribute; shape follows the attribute type (string, "
"number, date 'YYYY-MM-DD', bool, or list[str] for multi-select)."
),
)

@model_validator(mode="after")
def _value_actions_need_an_attribute(self) -> FacetAction:
"""The API 422s a value verb with no attribute_name; refuse it locally."""
value_verbs = (FacetActionType.set_value, FacetActionType.clear_value)
if self.action in value_verbs and not self.attribute_name:
raise ValueError(f"{self.action} needs an attribute_name")
return self

@classmethod
def classify(cls, content_type: ContentType | str) -> FacetAction:
"""Assign a content type, the batch form of `File.classify()`.

Args:
content_type: The content type to assign (object or path string).

Returns:
The action, unsent.
"""
return cls(
action=FacetActionType.classify, content_type_path=_path(content_type)
)

@classmethod
def unclassify(cls, content_type: ContentType | str) -> FacetAction:
"""Remove a content-type assignment, the batch form of `File.unclassify()`.

Args:
content_type: The content type to unassign (object or path string).

Returns:
The action, unsent.
"""
return cls(
action=FacetActionType.unclassify, content_type_path=_path(content_type)
)

@classmethod
def set_attribute(
cls, content_type: ContentType | str, name: str, value: Any
) -> FacetAction:
"""Set an attribute value, the batch form of `File.set_attribute()`.

Args:
content_type: The assigned content type (object or path string).
name: Attribute identifier (snake_case).
value: The value; shape depends on the attribute type (string, number,
date "YYYY-MM-DD", bool, or list[str] for multi-select).

Returns:
The action, unsent.
"""
return cls(
action=FacetActionType.set_value,
content_type_path=_path(content_type),
attribute_name=name,
value=value,
)

@classmethod
def clear_attribute(cls, content_type: ContentType | str, name: str) -> FacetAction:
"""Clear an attribute value, the batch form of `File.clear_attribute()`.

Args:
content_type: The assigned content type (object or path string).
name: Attribute identifier to clear.

Returns:
The action, unsent.
"""
return cls(
action=FacetActionType.clear_value,
content_type_path=_path(content_type),
attribute_name=name,
)

def _body(self) -> dict[str, Any]:
# The one place that knows the wire field names: the single-action methods
# on File post exactly this too, so single and batch can't drift.
body: dict[str, Any] = {
"action": self.action,
"content_type_path": self.content_type_path,
}
if self.attribute_name is not None:
body["attribute_name"] = self.attribute_name
if self.action == FacetActionType.set_value:
body["value"] = self.value # sent even when None, the server decides
return body


class FacetResult(BaseModel):
"""What one action in a `File.batch_facets()` returned, in request order.

Every result you receive succeeded: the endpoint fails fast, so a failing
action raises (see `LightOnAPIError.index`) and no results come back at all. A
returned list is therefore always complete and in order, so `results[i]` is the
outcome of `actions[i]`.
"""

model_config = ConfigDict(extra="ignore")

status: int = Field(
description=(
"Per-action status: 201 created, 200 already applied/updated, 204 for "
"the removals (unclassify, clear_attribute)."
)
)
data: dict[str, Any] | None = Field(
None,
description=(
"What the action returned, None for the 204 verbs. classify gives "
"{content_type_path, label}; set_attribute gives {name, value, "
"content_type_path, label}. Left a raw dict: it differs per verb, so "
"there is no one model to validate it into."
),
)


ContentType.model_rebuild() # resolve the self-referential `children` forward ref
Template.model_rebuild()
Loading
Loading