FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

bulk embeddings very slow · Issue #3753 · openai/openai-python · GitHub

Repository navigation

bulk embeddings very slow #3753

Description

Confirm this is an issue with the Python library and not an underlying OpenAI API

  • This is an issue with the Python library

Describe the bug

Symptons: Creating multiple embeddings with one api call is very slow (seconds vs miliseconds)
Where to look:
The bug is in https://github.com/openai/openai-python/blob/main/src/openai/lib/_parsing/_embeddings.py line 27
The code calls has_numpy() in a loop for every embedding. The function tries to import numpy, catches the exeption and returns a bool. If numpy isnt installed this is very slow.
Fix:
call has_numpy() either as global on instantiation, outside of the loop once or wrap has_numpy in functools.cache.

To Reproduce

  1. create embeddings for 1500 strings with one api call. Leave the encoding_format option empty
  2. meassure the time the api call takes
  3. do the same with encoding_format="base64", decode the entries with numpy manually and compare the time, it should be identical

Code snippets

OS

windows

Python version

Python v3.12.12

Library version

v3.5.0

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions


      Back | FazBrowse Home | New Git URL