Replacing Large Language Models with Jev Decision Models for Low-Latency Edge Service Orchestration
Abstract
Natural-language service requests can require a language-model decision before execution starts, consuming part of the request's latency budget. We integrate Jev's decision-oriented application programming interface (API) into edge service orchestration to reduce this overhead while retaining service completion. The integration extracts four to eight bounded intent fields and applies a shared validator, admission policy, and scheduler, accounting for decision waiting throughout the request timeline. We compare Jev, two self-hosted decision models, and three hosted large language models (LLMs) on 8,280 verified requests and on a live admission path with modeled execution and a real optical character recognition service. Across 33 test conditions, Jev reduces median decision latency by 22.7-64.5% relative to the fastest LLM. This latency barely moves with input size, contract width, or catalog size. On four-field contracts, Jev's API fees per correct decision are 59.7-80.9% lower at a cost of a few exact-match points, while wide contracts mark the limit of the substitution. Receiving the service catalog with each request, Jev names unseen services as accurately as known ones. On the live admission path, Jev keeps 0.91-0.95 of requests exact and on time at loads where the LLMs fall below 0.1. Since caching repeated descriptions gives the interpreters nearly the same latency, Jev's gain lies in fresh decisions. These results support decision-model substitution for latency-bound admission on bounded contracts.
Community
Can a decision model replace an LLM as the intent interpreter in edge service admission? We compare Jev with two self-hosted decision models and three hosted LLMs on 8,280 verified requests and on a live admission path with a real OCR service. Jev cuts median decision latency by 22.7–64.5% against the fastest LLM and keeps 0.91–0.95 of requests exact and on time at loads where the LLMs fall below 0.1. Wide eight-field contracts and cached repeat requests mark where its advantage ends.
Code: https://github.com/OniReimu/Edge-Computing-JEV
Data: https://huggingface.co/datasets/OniReimu/Edge-Computing-JEV
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Fast Intent-Driven Service Orchestration with Jev for 6G Edge Networks (2026)
- Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving (2026)
- Tessera: Demand-Driven KV Cache Management for Retrieval-Augmented LLM Serving (2026)
- Shared KV Caching for Replicated 27B Inference: Correctness Failures and Performance Boundaries (2026)
- Unified AI Gateway: A Framework for Joint Model Routing and KV Cache Management (2026)
- StreamDecisionBench: Evaluating Decisions in Force on Evolving Language Streams (2026)
- Herschel: Continuous Optimization of Production LLM Inference through On-Demand Profiling (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.22753 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper