GPU-களில் വേഗതയേറിയ എംബെഡിംഗുകൾ
Search, Computer മുതൽ ഞങ്ങളുടെ API Platform വരെയുള്ള Perplexity-യുടെ എല്ലാ കാര്യങ്ങൾക്കും വേഗതയേറിയതും കൃത്യവുമായ തിരച്ചിൽ വളരെ പ്രധാനമാണ്. തിരശീലയ്ക്ക് പിന്നിൽ, എംബെഡിംഗ്, റാങ്കിംഗ് മോഡലുകളാണ് കഠിനമായ ജോലികൾ ചെയ്യുന്നത്, നൽകിയിരിക്കുന്ന ഒരു ക്വറിക്കായി ഏറ്റവും പ്രസക്തമായ ഫലങ്ങൾ തിരിച്ചറിയാൻ ഇത് ഞങ്ങളെ സഹായിക്കുന്നു. state-of-the
Search, Computer മുതൽ ഞങ്ങളുടെ API Platform വരെയുള്ള Perplexity-യുടെ എല്ലാ കാര്യങ്ങൾക്കും വേഗതയേറിയതും കൃത്യവുമായ തിരച്ചിൽ വളരെ പ്രധാനമാണ്. തിരശീലയ്ക്ക് പിന്നിൽ, എംബെഡിംഗ്, റാങ്കിംഗ് മോഡലുകളാണ് കഠിനമായ ജോലികൾ ചെയ്യുന്നത്, നൽകിയിരിക്കുന്ന ഒരു ക്വറിക്കായി ഏറ്റവും പ്രസക്തമായ ഫലങ്ങൾ തിരിച്ചറിയാൻ ഇത് ഞങ്ങളുടെ സിസ്റ്റങ്ങളെ സഹായിക്കുന്നു. pplx-embed പോലുള്ള ഞങ്ങളുടെ സ്വന്തം മോഡലുകൾ പരിശീലിപ്പിക്കുകയും സെർവ് ചെയ്യുകയും ചെയ്തുകൊണ്ട് ഞങ്ങൾ അത്യാധുനിക ഗുണനിലവാരവും ലേറ്റൻസിയും കൈവരിക്കുന്നു.
ഈ പ്രത്യേക തരം മോഡലുകൾക്കായുള്ള Perplexity-യുടെ സെർവിംഗ് ഇൻഫ്രാസ്ട്രക്ചറിന്റെ തിരശ്ശീലയ്ക്ക് പിന്നിലെ ഒരു കാഴ്ച ഈ ലേഖനം അവതരിപ്പിക്കുന്നു. AI-നേറ്റീവ് തിരച്ചിലിന്റെ ഇൻഫറൻസ് ആവശ്യകതകളെ കാര്യക്ഷമമായി അഭിസംബോധന ചെയ്യുന്നതിനുള്ള ഞങ്ങളുടെ സാങ്കേതികവിദ്യകളെക്കുറിച്ച് ഞങ്ങൾ ചർച്ച ചെയ്യുന്നു, ഇത് ഞങ്ങളുടെ exabyte-scale search index-ന് ശക്തി പകരുമ്പോൾ തന്നെ മോഡലുകളുടെ വേഗത്തിലുള്ള പ്രോട്ടോടൈപ്പിംഗും വിലയിരുത്തലും സാധ്യമാക്കുന്നു. ഈ സാങ്കേതികവിദ്യകൾ സംയോജിമായി തിരയൽ ഗുണനിലവാരത്തിന്റെയും കാര്യക്ഷമതയുടെയും പരേട്ടോ അതിർവരമ്പ് വിപുലീകരിക്കുന്നു, ഏറ്റവും കുറഞ്ഞ ചിലവിലും ലേറ്റൻസിയിലും ഏജന്റുകൾക്കും ഉപയോക്താക്കൾക്കും സാധ്യമായ ഏറ്റവും മികച്ച ഫലങ്ങൾ നൽകാൻ ഞങ്ങളെ പ്രാപ്തരാക്കുന്നു.
Embeddings for Search
In a typical search setup, indexed documents are mapped to a high-dimensional vector സ്പേസ് using an embedding model and stored in a vector database. By embedding a query using the same model, similar documents can be located by finding the vectors closest to that of the query. This gives rise to two different traffic patterns for an inference engine to serve:
- ಬ್ಯಾಚ್ ಎಂಬೆಡಿಂಗ್: ಡೇಟಾಬೇಸ್ ಅನ್ನು ನಿರ್ಮಿಸುವಾಗ, ವಿಸ್ತರಿಸುವಾಗ ಅಥವಾ ಮರು-ಸೂಚ್ಯಂಕ ಮಾಡುವಾಗ, ವೆಕ್ಟರ್ ಸ್ಪೇಸ್ನಲ್ಲಿ ಬೃಹತ್ ದಾಖಲೆಗಳನ್ನು ಎಂಬೆಡ್ ಮಾಡಬೇಕು, ವೆಚ್ಚವನ್ನು ಕಡಿಮೆ ಮಾಡಲು ಥ್ರುಪುಟ್ ಅನ್ನು ಗರಿಷ್ಠಗೊಳಿಸಬೇಕು.
വെക്ടർ തിരച്ചിലിനെ തുടർന്ന്, വലിയ അളവിലുള്ള ഡോക്യുമെന്റുകൾ സ്കോർ ചെയ്യേണ്ടതുണ്ട്, ഇത് ത്രൂപുട്ടും ലേറ്റൻസിയും തമ്മിലുള്ള ബാലൻസ് നിലനിർത്തുന്നു.
- ഓൺലൈൻ എംബെഡിംഗ്: ഡാറ്റാബേസ് ക്വറി ചെയ്യുമ്പോൾ, ലുക്കപ്പുകൾക്കായി ഒരു ഹ്രസ്വ ക്വറി എംബെഡ് ചെയ്യേണ്ടതുണ്ട്, ഇത് ലേറ്റൻസി കുറയ്ക്കുന്നു.
ഉപയോഗ കേസുകളിലുടനീളം കഴിയത്ര പൊതുവായ ഘടകങ്ങൾ പ്രയോജനപ്പെടുത്തുന്നതിനായി ഞങ്ങൾ ഞങ്ങളുടെ ഇൻഫ്രാസ്ട്രക്ചർ നിർമ്മിച്ചു. ഞങ്ങൾ സാധാരണയായി എംബെഡിംഗുകൾ നിർമ്മിക്കാൻ ചെറിയ Transformer മോഡലുകൾ ഉപയോഗിക്കുന്നതിനാൽ, ഞങ്ങളുടെ LLM ഇൻഫറൻസ് കോഡുമായി നടപ്പിലാക്കലിന്റെ ഭൂരിഭാഗവും ഞങ്ങൾ പങ്കിടുന്നു: ബാച്ച് എംബെഡിംഗുകൾ കമ്പ്യൂട്ട്-ബൗണ്ട് പ്രെಫില്ലിന് സമാനമാണ്, അതേസമയം ഓൺലൈൻ എംബെഡിംഗുകൾ, പലപ്പോഴും കുറച്ച് ടോക്കണുകളിൽ പ്രവർത്തിക്കുന്നു, കമ്പ്യൂട്ടേഷണally മെമ്മറി-ബൗണ്ട് ഡീകോഡിന് സമാനമാണ്. അതിനാൽ എംബെഡിംഗ് മോഡലുകൾ സെർവ് ചെയ്യാൻ ഞങ്ങൾ ഞങ്ങളുടെ ഒപ്റ്റിമൈസ് ചെയ്ത പ്രെഫിൽ, ഡീകോഡ് കെർണലുകൾ വീണ്ടും ഉപയോഗിക്കുന്നു. ഫലമായി, ഓൺലൈൻ എംബെഡിംഗ് വർക്ക്ലോഡുകൾക്ക് കുറഞ്ഞ ലേറ്റൻസി നിലനിർത്തിക്കൊണ്ടുതന്നെ, കുറഞ്ഞ അധിക എഞ്ചിനീയറിംഗ് ജോലികളോടെ നമുക്ക് വലിയ ബാച്ച് ഇൻഫറൻസ് ത്രൂപുട്ട് കൈവരിക്കാൻ കഴിയും.
Tulips, Roses, and some Ivy
We expose inference through standardized APIs, both internally and externally through our API Platform. Under the hood, multiple services are involved in the processing of an embedding request:
- Ivy is a Rust HTTP gateway that Perplexity services call.
It handles the CPU-side work for requests such as JSON parsing, tokenization, input templating and batch splitting, translating requests to a custom gRPC protocol for downstream servers. This separation allows us to configure certain parameters around tokenization and input formatting without having to touch the heavier inference instances.
- Tulip is the inference server interface.
It is a gRPC server implemented with Rust, tokio, and `tonic`. Tulip receives gRPC inference requests, handling scheduling and batching. It then sends the batches to the ROSE engine, returning completed responses to clients.
- **ROSE** (Runtime-Optimized Serving Engine) മോഡൽ ഇൻഫറൻസ് നടപ്പിലാക്കുന്നു.
ഇത് പ്രധാനമായും Python-ൽ നിർവചിക്കപ്പെട്ടിരിക്കുന്നു, വൈവിധ്യമാർന്ന മോഡലുകൾക്കായി കെർണലുകൾ, ലെയറുകൾ, നിർവചനങ്ങൾ എന്നിവ നൽകുന്നു. ROSE മോഡലുകളിലൂടെയുള്ള ഫോർവേഡ് പാസ്സുകൾ നടപ്പിലാക്കുന്നു, കൂടാതെ എംബെഡിംഗുകൾക്കായി പ്രത്യേകമായ CUDA ഗ്രാഫ് മാനേജ്മെന്റും നൽകുന്നു. ഒരു ബാച്ച് സ്വീകരിക്കുകയും ആക്സിലറേറ്ററിൽ അത് നടപ്പിലാക്കുന്ന കണക്കുകൂട്ടലിന്റെ റഫറൻസ് തിരികെ നൽകുകയും ചെയ്യുന്ന ഒരു step() ഫംഗ്ഷൻ വഴി ഇത് Tulip-ലേക്ക് ബന്ധിപ്പിച്ചിരിക്കുന്നു.

Paying Attention Beyond the Kernel
Transformer-അധിഷ്ഠിത മോഡലുകളും അടിസ്ഥാന Hopper/Blackwell ആർക്കിടെക്ചറുകളും പക്വത പ്രാഴിച്ച സാങ്കേതികവിദ്യകളാണ്, അതിനാൽ GPU വശത്ത് എംബെഡിംഗ് ഇൻഫറൻസ് വിവിധ ഇൻഫറൻസ് എഞ്ചിനുകളിലുടനീളം ഒപ്റ്റിമൽ ആയ ഒരു നടപ്പിലാക്കലിലേക്ക് എത്തിച്ചേർന്നിട്ടുണ്ട്. എങ്കിലും, ക്ലയന്റിന് മോഡലുകൾ എൻഡ്-ടു-എൻഡ് ലഭ്യമാക്കുന്ന റൺടൈമുകളിലും ഹാർനെസുകളിലും മെച്ചപ്പെടുത്തലുകൾക്കുള്ള അധിക അവസരങ്ങൾ ഞങ്ങൾ കണ്ടെത്തി. പ്രത്യേകിച്ചും, CUDA ഗ്രാഫുകൾ ശ്രദ്ധാപൂർവ്വം കൈകാര്യം ചെയ്യുന്നതിലൂടെയും നേറ്റീവ് Rust എഞ്ചിനിൽ GPU-വശത്തുള്ള ഫലം അസിൻക്രണസ് ആയി ട്രാക്ക് ചെയ്യുന്നതിനായി ഒരു LazyTensor അമ്രാക്ഷൻ നിർമ്മിക്കുന്നതിലൂടെയും ഞങ്ങൾക്ക് ലേറ്റൻസി മെച്ചപ്പെടുത്താൻ കഴിയുമെന്ന് ഞങ്ങൾ കണ്ടെത്തി. ROSE-ന്റെ മോഡൽ ഇംപ്ലിമെന്റേഷനുകളുമായി ഇത് ഫലപ്രദമായി സംവദിക്കാൻ കഴിയുന്ന തരത്തിൽ ഞങ്ങൾ ഈ ഫീച്ചറുകൾ Tulip-ൽ നടപ്പിലാക്കി.
Tulip
We designed Tulip to be as lightweight of an interface over our model serving as possible. It handles incoming requests in Tokio async tasks, maintaining a pool of requests it tracks and schedules batches from to dispatch to the accelerator. The scheduling mechanism in Tulip is very simple: requests accumulate while Tulip is dispatching work or waiting for results. From the accumulated requests, sequences are picked on a first-come, first-served basis to be run through the model.
The simple scheduling mechanism is motivated by an observation on model performance. For small embedding models, at the sequence lengths we serve for, we noticed that the linear cost of dense layers is dominant over the quadratic cost of attention. Thus, latency is mostly proportional to the number of tokens, not the number of sequences. Consequently, once a batch is large enough to saturate the GPU, which is around 512 tokens on a model under one billion parameters, packing more sequences into it does not improve efficiency.
To effectively interface with the model, Tulip relies on CUDA graphs and lazy result tracking to overlap GPU and CPU work and fully utilize the available resources.
CUDA ഗ്രാഫ് മാനേജ്മെന്റ്
एक मॉडल के फॉरवर्ड पास को चलाने में सीपीयू-साइड और जीपीयू-साइड दोनों कार्य शामिल होते हैं। सीपीयू उचित मापदंडों के साथ बैचों को शेड्यूल करने और कर्नेल लॉन्च करने के लिए ज़िम्मेदार है, जबकि जीपीयू प्रासंगिक मैट्रिक्स गुणन, ध्यान, आदर्श या सक्रियण कर्नेल निष्पादित करता है। प्रशिक्षण और पुनः अनुक्रमण जैसे उच्च-थ्रूपुट कार्यभार के लिए, सीपीयू-साइड ओवरहेड नगण्य हैं क्योंकि बैच आकार और जीपीयू-साइड विलंबता दोनों बड़े हैं। हालाँकि, छोटे बैच आकारों पर, सीपीयू-साइड का काम जीपीयू-साइड के काम से अधिक हो सकता है।

To mitigate overheads, instead of launching independent kernels, a CUDA graph can be built to capture the metadata required to launch all the kernels of a forward pass with a single call to the CUDA driver. This eliminates the need to re-run expensive Python and PyTorch code for the configurations CUDA graphs can be captured for.
ഓരോ മോഡലിലുടനീളം, GPU എക്സിക്യൂഷൻ CPU-വശത്തുള്ള കെർണൽ ലോഞ്ചിനേക്കാൾ ചെലവേറിയതാകുന്ന ടോക്കണുകളുടെ ഏറ്റവും കുറഞ്ഞ എണ്ണം നിർണ്ണയിച്ചുകൊണ്ട് ഞങ്ങൾ ഒരു ഇൻഫ്ലക്ഷൻ പോയിന്റ് ട്രാക്ക് ചെയ്യുന്നു. എംബെഡിംഗ് മോഡലുകൾ ചെറുതായതിനാൽ, ആയിരക്കണക്കിന് ടോക്കണുകളുടെയും പത്തانுகു സീക്വൻസുകളുടെയും ബാച്ചുകളിലാണ് ഈ ഇൻഫ്ലക്ഷൻ പോയിന്റ് വരുന്നതെന്ന് ഞങ്ങൾ നിരീക്ഷിക്കുന്നു. ചില അറ്റൻഷൻ നടപ്പിലാക്കലുകൾ കെർണൽ ലോഞ്ചുകൾ കോൺഫിഗർ ചെയ്യാൻ ഡൈനാമിക് ഹോസ്റ്റ്-സൈഡ് ഇൻപുട്ടുകളെ ആശ്രയിക്കുന്നു, ഇത് പൂർണ്ണ-മോഡൽ പ്രെഫിൽ/ഡെൻസ് CUDA ഗ്രാഫുകളെ തടയുന്നു. ഞങ്ങളുടെ ഇൻഫറൻസ് എഞ്ചിനിൽ അവ പ്രവർത്തനക്ഷമമാക്കാൻ പ്രസക്തമായ കെർണലുകളിലേക്ക് ഞങ്ങൾ മാറ്റങ്ങൾ അപ്സ്ട്രീം ചെയ്തു.
To address overheads, we build whole-model CUDA graphs for all embedding models and overlap CPU work with GPU work. Since CUDA graphs minimize the CPU-side overheads, once a graph is launched, we have free time to kick off and enqueue the execution of the next batch whenever it is available. The results of the pending batch are tracked with a LazyTensor, which allows an async task in Rust to block until the previous batch finishes execution. CUDA graphs help low-latency serving by ensuring we are not held back by the cost of kernel launches and facilitate improved scheduling in the high-throughput case as they free up the CPU to do work on the next batch sooner.

CUDA graphs must be captured for each distinct configuration, which for embeddings means a graph per sequence count and token count combination. Since this grid is expansive, we pad token counts to buckets that are multiples of 64 or 256. This still results in thousands of graphs that might take multiple minutes to capture for a typical model. The cost of capture comes from two sources: an eager forward pass that must be executed to compile kernels and set up buffers for various kernels that need them, followed by the capture run which re-executes Python code.
We mitigate startup costs by capturing CUDA graphs lazily as the engine serves. We keep track of each configuration and ensure that it goes through an eager warmup run before triggering graph capture and replay on the second hit. All subsequent executions of the same graph configuration then go through CUDA graph replay. Lazy graph capture has an impact on p99 latencies during startup; however, it is valuable in spreading multiple minutes of eager work across multiple hours. Quicker startup times allow us to better scale and manage embedding deployments.
ലസി ടെൻസറുകൾ (Lazy Tensors)
Through CUDA, GPU work is asynchronous. Since launching a kernel asynchronously enqueues it on a stream, host code must explicitly synchronize to read out the resulting vectors. To facilitate a higher degree of parallelism and to be able to kick off future batches while waiting for the previous one to complete on the device, we rely on a LazyTensor abstraction to track values.
The LazyTensor tracks a host buffer in page-locked memory and a cudaMemcpyAsync operation via an event copying data from the device. It is kicked off after the launch of the forward pass on the same stream. Since the copy operation must wait for all prior kernels on the stream to execute, the associated event tracks both the completion of the forward pass and the availability of the result on the CPU.

GPU, CPU ജോലികൾ ഓവർലാപ്പ് ചെയ്യുന്നതിനായി ഞങ്ങൾ ഞങ്ങളുടെ ROSE എൻകോഡർ എഞ്ചിനിൽ LazyTensor-കൾ പ്രയോജനപ്പെടുത്തുന്നു. ഓരോ step() കോളും CUDA ഗ്രാഫ് പ്രവർത്തിപ്പിച്ച് അത് തീരുന്നതുവരെ കാത്തിരിക്കുന്നതിന് പകരം, step() അതിന്റെ ഫലം അസിൻക്രണസ് ആയി ട്രാക്ക് ചെയ്യാൻ ഒരു LazyTensor തിരികെ നൽകുന്നു. CUDA ഗ്രാഫുകളുമായി സംയോജിപ്പിച്ച്, കുറഞ്ഞ ലേറ്റൻസിയും മെച്ചപ്പെട്ട ത്രൂപുട്ടും കൈവരിക്കാൻ ഇത് ഞങ്ങളെ സഹായിക്കുന്നു.

ROSE
LLM സെർവിംഗിനായി ഞങ്ങൾ യഥാർത്ഥത്തിൽ നിർമ്മിച്ച ഞങ്ങളുടെ ROSE എഞ്ചിൻ, എംബെഡിംഗ് മോഡലുകളുടെ എക്സിക്യൂഷൻ കൂടി കൈകാര്യം ചെയ്യുന്നതിനായി ഞങ്ങൾ അനുയോജ്യമാക്കി. എംബെഡിംഗ് മോഡലുകൾക്കുള്ള പിന്തുണയ്ക്കേണ്ട പരിശ്രമം കുറയ്ക്കാൻ, ROSE LLM-களும் എംബെഡിംഗുകളും തമ്മിൽ കോഡ് ആക്രമണോത്സുകമായി വീണ്ടും ഉപയോഗിക്കുന്നു. ഉദാഹരണത്തിന്, pplx-embed സെർവിംഗും Qwen3.5 LLM ഡീകോഡിംഗും എല്ലാം ഒരേ കെർണലുകളിലൂടെയാണ് കടന്നുപോകുന്നത്. പ്രോട്ടോടൈപ്പിംഗ്, വിലയിരുത്തൽ, പ്രൊഡക്ഷൻ ഇൻഫറൻസ് എന്നിവയ്ക്കായി ഒരു LLM-ൽ നിന്ന് യഥാർത്ഥത്തിൽ ഫൈൻ-ട്യൂൺ ചെയ്ത ഒരു എംബെഡിംഗ് മോഡൽ എളുപ്പത്തിൽ സെർവ് ചെയ്യാൻ ഈ പങ്കിടൽ ഞങ്ങളെ അനുവദിക്കുന്നു.
For dense layers, embedding and LLM inference are identical since token vectors are processed independently. In attention layers, differences are handled by adding support for ragged inputs, alongside the paged prefill and decode setups required by LLMs. When serving an embedding model, we do not instantiate a KV cache and dispatch to variations of attention kernels which support the ragged format to avoid padding. The supporting conversion and calibration routines are also shared with the LLMs.
Ivy
ഞങ്ങളുടെ ഇൻഫറൻസ് HTTP പ്രോക്സി ലെയറായ Ivy-യും പ്രകടനത്തിൽ പ്രധാനപ്പെട്ട ഒരു പങ്ക് വഹിക്കുന്നു. പ്രൊഡക്ഷനിൽ റിക്വസ്റ്റ് പേലോഡുകൾ വ്യത്യാസപ്പെടുന്നതിനാൽ, വ്യക്തിഗത റിക്വസ്റ്റുകൾ വ്യക്തിഗത റെപ്ലിക്കകളിലേക്ക് റൂട്ട് ചെയ്യുന്നത് ലോഡ് അസമത്വത്തിന് കാരണമാകും. Ivy വലിയ ബാച്ച് റിക്വസ്റ്റുകളെ കഷ്ണങ്ങളായി വിഭജിക്കുകയും റെപ്ലിക്കകൾക്കിടയിൽ അവയുടെ ലോഡ് ബാലൻസ് ചെയ്യുകയും ചെയ്യുന്നു, ഇത് ഉപയോഗം മെച്ചപ്പെടുത്തുകയും ലേറ്റൻസി സുഗമമാക്കുകയും ചെയ്യുന്നു. Ivy-യിൽ പൂർണ്ണമായി സജീവമാക്കിയ ഞങ്ങളുടെ ഇൻ-ഹൗസ് യൂണിഗ്രാം ടോക്കണൈസേഷനിലെ പുതിയ ജോലി, ഓഫ്-The-Shelf ടോക്കണൈസറുകളെ അപേക്ഷിച്ച് ലേറ്റൻസി ഗണ്യമായി മെച്ചപ്പെടുത്തുന്നു.
...but the Kernels Still Matter
ROSE supports a variety of attention backends. Different kernels may be suited to specific problem sizes. Over time, we integrated FlashInfer 2, FlashInfer 3 and FlashAttention 4 kernels to implement ragged attention.

പൊതുവേ, FlashAttention 4 വേഗതയേറിയതാണെന്ന് ഞങ്ങൾ കാണുന്നു. എന്നിരുന്നാലും, വളരെ ദീർഘമായ സീക്വൻസ് നീളത്തിൽ Qwen-അധിഷ്ഠിത മോഡലുകളിൽ FlashInfer 3 ഇതിനേക്കാൾ മികച്ച പ്രകടനം കാഴ്ചവയ്ക്കുന്നു. പെർഫോമൻസും ട്യൂണിംഗും അറ്റൻഷൻ ഹെഡുകളുടെ എണ്ണത്തിനും അളവിനും അനുസരിച്ച് വ്യത്യാസപ്പെടാമെന്നതിനാൽ, ഞങ്ങൾ ഒന്നിലധികം കോൺഫിഗറേഷനുകൾക്കുള്ള പിന്തുണ നിലനിർത്തുകയും സെർവ് ചെയ്യുമ്പോൾ ഓരോ കേസിനും അനുയോജ്യമായ തീരുമാനം എടുക്കുകയും ചെയ്യുന്നു.
Benchmarks
मूल्यांकन डेटासेट से प्राप्त वास्तविक मॉडल वेट और इनपुट पर BF16 परिशुद्धता पर अनुमान चलाते हुए, हम vLLM v0.22.0 के विरुद्ध बेंचमार्क करते हैं। सभी टाइमिंग रन से पहले वार्मअप रन किए गए थे, जिन्होंने सत्यापित किया कि कोसाइन समानता में विचलन 0.1% के भीतर है।
കുറഞ്ഞ ലേറ്റൻസിയുള്ള എംബെഡിംഗുകൾ (p50 / p90 / p99 / പരമാവധി ms)
We report runtimes for pre-tokenized request batch size 1, fully sequential requests, sequence lengths of 128, 512, and 4096 tokens.

Low-Latency Scoring (p50 / p90 / p99 / max ms)
Pre-tokenized request batch sizes 5, 25, and 50, sequence length of 512 tokens.

ഉയർന്ന ത്രൂപുട്ടുള്ള എംബെഡിംഗുകൾ (emb/s)
റിക്വസ്റ്റ് ബാച്ച് സൈസ് 100, റിക്വസ്റ്റുകൾ സമർപ്പിക്കുന്ന നാല് ഒരേസമയം പ്രവർത്തിക്കുന്ന പ്രോസസ്സുകൾ, 512, 1024, 4096 ടോക്കണുകളുടെ സീക്വൻസ് നീളങ്ങൾ.

High-Concurrency Embeddings (p50 / p90 / p99 / max ms)
സീക്വൻസ് നീളം 512, ബാച്ച് സൈസ് 1, എന്നാൽ ഞങ്ങൾ 1, 2, 4, 8, 16 ഒരേസമയം നടക്കുന്ന റിക്വസ്റ്റുകൾ അയക്കുന്നു. ഈ ബെഞ്ച്മാർക്കിൽ Ivy വഴിയുള്ള ടോക്കണൈസേഷൻ ചെലവുകളും, Ivy-യും Tulip-ഉം തമ്മിലുള്ള നെറ്റ്വർക്കിംഗ് ഓവർഹെഡും ഉൾപ്പെടുന്നു.

Conclusion and Future Work
Ivy, Tulip, ROSE എന്നിവ ഉൾക്കൊള്ളുന്ന സെർവിംഗ് ഇൻഫ്രാസ്ട്രക്ചർ Perplexity-ക്കായി കുറഞ്ഞ ലേറ്റൻസിയിലും മികച്ച ത്രൂപുട്ടിലും എംബെഡിംഗുകൾ സെർവ് ചെയ്യാൻ ഞങ്ങളെ അനുവദിക്കുന്നു, ഇത് മുൻകൂട്ടി തയ്യാറാക്കിയ സൊല്യൂഷനുകളെ അപേക്ഷിച്ച് കുറഞ്ഞ ചെലവിൽ കൂടുതൽ കൃത്യമായ തിരച്ചിലിന് കാരണമാകുന്നു.
ನಿರ್ದಿಷ್ಟ ಮಾದರಿಗಳ ಮೇಲೆ ಗಮನಹರಿಸುವ ಮೂಲಕ ಮತ್ತು ಸಂಪೂರ್ಣ ಸ್ಟಾಕ್ನ ಮಾಲೀಕತ್ವವನ್ನು ತೆಗೆದುಕೊಳ್ಳುವ ಮೂಲಕ, ಕಾರ್ಯಕ್ಷಮತೆ ಮತ್ತು ನಮ್ಯತೆಯ ನಡುವೆ ಪರಿಣಾಮಕಾರಿ ಸಮತೋಲನವನ್ನು ಸಾಧಿಸಲು ನಮಗೆ ಅಗತ್ಯವಿರುವ ಸ್ವಾತಂತ್ರ್ಯ ದೊರೆಯುತ್ತದೆ. ಹೆಚ್ಚು ಮರುಬಳಕೆ ಮಾಡಬಹುದಾದ ಮತ್ತು ಕಾರ್ಯಕ್ಷಮತೆಯ Rust ಪ್ರಿಮಿಟಿವ್ಗಳನ್ನು ಹೆಚ್ಚು ಜೆನೆರಿಕ್ Python ಮಾಡಲಿಂಗ್ ಕೋಡ್ನೊಂದಿಗೆ ಮಿಶ್ರಣ ಮಾಡಲಾಗುತ್ತದೆ. vLLM, SGLang, ಮತ್ತು TokenSpeed ನಂತಹ ಅನೇಕ ಓಪನ್ ಸೋರ್ಸ್ ಇನ್ಫರೆನ್ಸ್ ಎಂಜಿನ್ಗಳು Rust ಮತ್ತು C++ ನಂತಹ ಭಾಷೆಗಳನ್ನು ತಮ್ಮ ಸ್ಟಾಕ್ನಲ್ಲಿ ಸಂಯೋಜಿಸುತ್ತಿವೆ. ಕಳೆದ ಎರಡು ವರ್ಷಗಳಲ್ಲಿ ನಾವು Rust ನಲ್ಲಿ ಹೂಡಿಕೆ ಮಾಡಿದ್ದೇವೆ ಮತ್ತು ಕಾರ್ಯಕ್ಷಮತೆ ಹಾಗೂ ನಿರ್ವಹಣೆಯಲ್ಲಿ ದೊಡ್ಡ ಲಾಭವನ್ನು ಪಡೆದಿದ್ದೇವೆ. ನಮ್ಮ LLM ಸರ್ವಿಂಗ್ ಸ್ಟಾಕ್ನೊಂದಿಗೆ ಹೆಚ್ಚಿನ ಎಂಬೆಡಿಂಗ್ ಅನುಷ್ಠಾನವನ್ನು ಹಂಚಿಕೊಳ್ಳುವ ಮೂಲಕ, ಎಂಬೆಡಿಂಗ್ ಮಾದರಿಗಳ ನಿರ್ವಹಣೆಗಾಗಿ ಗಮನಾರ್ಹವಾದ ಎಂಜಿನಿಯರಿಂಗ್ ಪ್ರಯತ್ನವನ್ನು ব্যয়ಿಸದೆ, ನಾವು ಥ್ರುಪುಟ್ನಲ್ಲಿಯೂ ಲಾಭವನ್ನು ಗಳಿಸುತ್ತೇವೆ.
As models evolve, we will continue to improve each layer of our stack to reduce both CPU-bound and GPU-bound latencies. Our custom gRPC-based protocols within Ivy and Tulip allow us to tweak communication to reduce network latencies, while ROSE provides a foundation to improve computational throughput. Additionally, as support for free-threaded Python grows throughout the ecosystem, we will be able to further improve Python-Rust interoperability to reduce overheads.