Skip to content

请问,web server功能,运行了之后,发现调用/v1/chat/completions接口的回复不完整,回复内容很短,请问这个问题在哪里可以调整呢 #421

Description

@12lxr

Is your feature request related to a problem? Please describe.
A clear and concise description of what the problem is. Ex. I'm always frustrated when [...]

Describe the solution you'd like
A clear and concise description of what you want to happen.

Describe alternatives you've considered
A clear and concise description of any alternative solutions or features you've considered.

Additional context
Add any other context or screenshots about the feature request here.

Activity

  1. gjmulder commented on Jun 26, 2023

    @gjmulder
    Contributor

    Sorry, but you will have to ask your question in English.

    You would want to choose a model that is designed for chat, and ensure your prompt is similar in format to how that specific model was trained.

    Once you are using a model trained for chat, try initialising the model with a n_ctx=1024. When prompting using the defaults you will get max_tokens=128. Try setting this to a value to say half your n_ctx.

  2. Repository owner locked and limited conversation to collaborators on Jun 26, 2023
  3. converted this issue into a discussion #425 on Jun 26, 2023
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    qualityQuality of model output

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions