< Back to situations

We’ll email you as it develops, and you can follow the whole thread from day one.

[SITUATION] · [ACTIVE]

2 clusters · 2 sources · 20 days · First seen · Last updated

Categories: TECHNOLOGY

AI routing and on‑premise model deployment

Entities: Ollama · AT&T Inc. · insurance industry · Tools4AI · Gemma 4

Overview

Early in July, AI teams began experimenting with routing policies and semantic caching aimed at lowering the cost of running large language models. The approach involved directing inference requests to the most appropriate model based on task complexity, and using cache mechanisms to reuse previously computed results.

By late July, AT&T unveiled an AI routing framework that dynamically selects among several open‑source models—ranging from smaller, inexpensive options to larger, high‑performing ones—achieving up to a 90 % reduction in inference expenses while maintaining latency and response quality. The system integrates models such as Google’s Gemma 4, OHI‑4, OSS‑120B, and a telecom‑specific OTel 2.0 model.

Concurrently, the pure‑Java agentic AI platform Tools4AI, paired with the Ollama runtime, enabled enterprises in regulated sectors like insurance to run AI agents on‑premise. A demo showcased a claims‑triage agent that ingests free‑text incident reports, routes them to appropriate business actions, extracts structured data, and enforces human oversight, ensuring sensitive data never leaves the organization.

Timeline

  1. 1 day ago

    [TECHNOLOGY] 2 sources
    AT&T AI Routing Cuts Inference Costs; Tools4AI Brings On‑Premise Agents to Insurance

    AT&T’s AI routing framework slashes inference costs by up to 90% using open‑source models, while Tools4AI with Ollama enables on‑premise AI agents for insurance claim processing, keeping data inside the firm.

  2. 21 days ago

    [TECHNOLOGY] 2 sources
    AI Teams Implement Routing Policies and Semantic Caching to Cut Model Costs

    AI teams are urged to adopt routing policies and semantic caching to document model choices and reuse similar queries, lowering costs and improving auditability.

Sources

dev.to · opensourceforu.com