Skip to content

[OMEGA-427] Use tools API to pass the list of tools to the LLM - #358

Open
vsbogd wants to merge 25 commits into
singnet:mainfrom
vsbogd:use-tools-api
Open

vsbogd wants to merge 25 commits into
singnet:mainfrom
vsbogd:use-tools-api

Conversation

@vsbogd

@vsbogd vsbogd commented Sep 21, 2026 •

Copy link
Copy Markdown
Member

Description

This PR is part of #349 which doesn't include loop and memory changes. It delivers the most massive code change: using tools API to get the list of tool calls from LLM. It obsoletes helper.balance_parenthesis function.

How Has This Been Tested?

This PR doesn't introduce any code logic change just migrates existing logic to the new API. Regression testing should be enough. This PR is checked using automatic regression testing and manual smoke check with each provider.

Checklist

  • PR contains autogenerated code
  • Self-review completed
  • Test scenarios above are passed with the version of the code from PR

Move provider specific logic into prepare_args method mostly. The only
exceptions are OpenAI because it uses completely different method to
call the API and TestMock which doesn't require most of the things.
Change the LLM call API to pass the list of tools to call and receive
the list of tool calls. Adapt loop.metta to this change. Fix unit tests.
The historical information is still returned via HISTORY paragraph of
the system prompt.
@vsbogd
vsbogd marked this pull request as ready for review September 23, 2026 09:03
@vsbogd vsbogd changed the title [OMEGA-374] Use tools API to pass the list of tools to the LLM [OMEGA-427] Use tools API to pass the list of tools to the LLM Sep 24, 2026

@paul-v-snet paul-v-snet left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Other non-critical problems:

Also, most of the issues found during QA in #349 still apply to this PR, so they should be addressed here as well.

Comment thread providers/lib_llm_ext.py Outdated
Comment thread providers/openai.py Outdated
Comment thread providers/lib_llm_ext.py Outdated
Comment thread src/providers.metta
Comment thread src/providers.py Outdated
Comment thread src/providers.py
Comment thread src/providers.py
Comment thread plugins/workflow/instructions/research-workflow/skill.metta Outdated
Comment thread src/loop.metta Outdated
@vsbogd

vsbogd commented Sep 28, 2026

Copy link
Copy Markdown
Member Author

Fixed in 1d31d57

paul-v-snet
paul-v-snet previously approved these changes Sep 28, 2026
@vsbogd

vsbogd commented Sep 28, 2026

Copy link
Copy Markdown
Member Author

@TossSky please pay attention on the following three commits:

  • badf680 - integration test is fixed after fixing llmToolCallToSExpr
  • 80ab9b6 - integration test is fixed after fixing llmToolCallToSExpr
  • 027b3bc temporarily make test_llm_budget.py optional

@jazzbox35 jazzbox35 mentioned this pull request Sep 29, 2026
1 of 3 tasks

@paul-v-snet paul-v-snet left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Discovered issues:

  1. addToHistory is still called with a single argument, while it expects 3. I'm not sure what will happen, so I just highlighted it.
  2. Tool calling without strict mode may cause issues. I suggested changes that fix this.

Comment thread plugins/workflow/workflow.metta Outdated
Comment thread src/providers.py
Comment thread providers/openai.py Outdated
Comment thread providers/lib_llm_ext.py Outdated
Comment thread src/providers.py
@@ -1,9 +1,204 @@
import config
import logging

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
import logging
import logging
import json

Comment thread src/providers.py

def set_tool(self, tool: LLMTool):
self.tool = tool

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
def stringify_arguments(self):
self.arguments = {
name: value if isinstance(value, str) else json.dumps(value, ensure_ascii=False)
for name, value in self.arguments.items()
}

Comment thread src/providers.py

def _validate_response(request: LLMRequest, response: LLMResponse) -> LLMResponse:
for call in response.calls:
if call.is_error():

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
if call.is_error():
call.stringify_arguments()
if call.is_error():

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My opinion such conversion can hide issues. I don't fully understand whether it is possible for API implementor send non-string argument if type of the parameter is "string" and "strict" is True. And if it is possible we probably should consider it as a bug on the LLM provider side. I suggest wait for such issue to be reproduced before fixing it.

@TossSky

TossSky commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator

@vsbogd @paul-v-snet @jazzbox35
Tested: image built from c193ab3, Omega version=v0.1.19-166-gc193ab3. The CI suites ran with the agent started as in CI, the WebSocket, Slack and Telegram mock suites on their own channels, and the live IRC suite once per provider: Anthropic claude-opus-4-8, OpenAI gpt-5.5, OpenRouter z-ai/glm-5.2, ASICloud minimax/minimax-m3. The ASIOne key I have is rejected with 401, so ASIOne is not covered. I reran the failed live tests on main at 6657798, which has the same product code as ee0618a.

What I checked:

  • the tools go in the tools field, the system prompt no longer lists skills, and the request carries max_tokens and reasoning_mode
  • each request carries the previous iteration's calls and their results under the same ids
  • Anthropic, OpenAI, OpenRouter and ASICloud accept the strict schemas, tool_choice: "required" and tools without parameters
  • tests/mettatest.sh: 6 of 6 OK
  • tests/pytest.sh: 78 passed, 1 skipped
  • pytest @run_mandatory: 135 passed, 1 failed; test_git_local_commit_mock then passed three reruns
  • @run_mandatory also collects on Python 3.10: 136 tests
  • import_knowledge/: 21 passed
  • mock_websocket/: 13 passed
  • mock_slack/: 29 passed, 4 failed, 1 skipped; mock_telegram/: 19 passed, 4 failed, 1 skipped (item 6)
  • live IRC suite: Anthropic 33 passed / 3 failed, OpenRouter 32 / 4, OpenAI 27 / 9, ASICloud 25 / 11
  • a read-file of a missing or empty path no longer stops the agent
  • an unknown tool or a missing parameter comes back to the model as an error result, and the other calls of the response still run
  • quotes, backslashes, a literal \n, tabs, CR LF, non-ASCII text and a NUL byte reach write-file unchanged
  • delegation results reach the model through HISTORY
  • the workflow plugin no longer writes to the history; after a load the user gets Loaded workflow: ... and the next request lists the workflow's tools, and a second unload answers No active workflow is loaded

On every provider test_create_script.py and test_complex_weather_flow.py fail for the reasons described in #349. I leave test_git_push_to_remote.py out, since the token I had for the test repository has expired. test_edit_add_timestamp.py fails on OpenRouter on main as well.

  1. A call that gets no result makes Anthropic and OpenAI reject the next request, and the agent stops answering. When a skill fails in MeTTa instead of returning a value, for example read-file of a missing file or metta with a match that finds nothing, the call is already in the assistant message (loop.metta:115) but its tool message is never added (loop.metta:121-125). Anthropic answers An assistant message with 'tool_calls' must be followed by tool messages responding to each 'tool_call_id', OpenAI answers No tool output found for function call fc_.... The model never sees an outcome and repeats the call, so every second request is rejected until the iterations run out:

                                              requests  rejected  reply to the user
    Anthropic, main   metta, empty match           45         0  "It returned empty []"
    Anthropic, #358   metta, empty match           50        25  none
    OpenAI,    #358   metta, empty match           33         8  none
    Anthropic, #358   read-file of a missing file  50        24  none
    

    On main the missing-file question goes unanswered too, but without rejected requests. The regular live suite on Anthropic hit the same rejection three times. A mocked answer reproduces it: [("read-file", {"filename": "/tmp/missing.txt"}), ("send", {"content": "after"})] gives a next request with two calls and one result. A non-deterministic metta result breaks the pairing the other way: (superpose (1 2 3)) yields three results under one call id, which Anthropic rejects as well ('tool_call_id' of 'toolu_...' not found in 'tool_calls' of previous message), 25 of 50 requests.

  2. With tool_choice: "required" (lib_llm_ext.py:203, openai.py:86) and no no-op tool (nop is defined but not registered), the model has to call something on every iteration, and the prompt tells it to query first (prompt.txt:6). After a single question on Anthropic the agent made 104 query calls over 53 iterations, against 22 over 23 on main. None of them repeats a query from the previous iteration, whose results are in the request, while 27 of the 104 repeat a query from two iterations back, whose results are already gone. On OpenAI I reran the six live tests that failed (test_memory_history, test_run_error_script, test_run_repeated, test_search_invalid, test_skill_episodes, test_skill_metta):

           passed  requests  query  send   time
    main     3/6      438     383     5   566 s
    #358     1/6      765     607     4  1049 s
    

    On Anthropic a simple question still gets its answer in 2 to 8 seconds, as on main.

  3. On ASICloud 61 of 193 responses came without any tool call despite tool_choice: "required", and the loop drops them. In a direct request the model wrote the call as text instead, [TOOL_CALL]\n{"tool":"version","args":{}}\n[/TOOL_CALL, with finish_reason stop. Of the other eight ASICloud live failures, five pass on main (test_edit_add_timestamp, test_edit_append_line, test_search_basic, test_search_weather, test_skill_metta). ASICloud is the launcher's default provider (scripts/omega:569).

  4. The fallback call id {response_id}#{uuid} (lib_llm_ext.py:71) is still rejected by two providers. With max_tokens 24 Anthropic returned finish_reason length, the fallback got the id msg_011CfXndPVaGDxcDvUYYQix7#36bcc9f7be2e458fad05f2cbc4130d54, and the next request failed with tool_use.id: String should match pattern '^[a-zA-Z0-9_-]+$'. OpenAI rejects the same construction by length, 88 characters against the limit of 64. OpenRouter and ASICloud accept it.

  5. Loading the same workflow twice without unloading duplicates its tools: after two workflow-load-instructions calls for test-workflow the request lists test-skill three times (workflow.metta:80-86, skills.metta:31). Anthropic rejects any request with a repeated tool name (Duplicate tool or function name: send), and the copies stay until workflow-unload-instructions. In the live run Claude unloaded before loading again, so this only happens with a model that loads twice.

  6. Phase 2 of CI runs nothing. unit/test_llm_budget.py still tests the old chat(str) contract and stubs providers without LLMToolCall, so its collection fails with a NameError at lib_llm_ext.py:66 and pytest stops the whole phase (job log): collected 6 items / 1 error, Interrupted: 1 error during collection. continue-on-error keeps the job green. Three of the other six optional tests pass when I run them one at a time. test_complex_weather_flow_mock.py:62 and test_run_create_dirs_mock.py:49 fail because their answers still carry a literal \\n, which now stays a backslash and an n. The change made in badf680 and 80ab9b6 was not applied to them, nor to the Slack and Telegram copies of create_script, edit_append_line, run_create_dirs and complex_weather_flow, which are the four failures in each of those suites. test_git_push_to_remote_mock runs into the 5-second shell limit on git clone ... && git push and gets timeout_error, the same as on [OMEGA-374] Omega V2: using tools API fields to pass the list of tools to LLM #349.

  7. A tool argument that is not a string stops the agent. For an int, float, bool, None, list or dict value llmToolCallToSExpr raises AttributeError: 'int' object has no attribute 'translate' (providers.py:329), which is not caught (loop.metta:109), and PeTTa exits with code 2. None of the four providers sent a non-string when asked for the integer 42: Anthropic, OpenAI and ASICloud sent "42", and GLM refused.

  8. tests/test_src_providers.py defines test_validate_repsonse_no_tool twice, so pytest collects three tests. The first one, where a valid call passes validation, never runs, and it would fail on the undefined expected and on with_tool, which LLMToolCall does not have.

I fixed items 1, 4, 5, 6, 7 and 8 and opened vsbogd/Omega#9 into this branch. Items 2 and 3 need a decision about the loop.

Verdict: FAIL

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants