Skip to main content Signal blog Official Microsoft Blog Command Line Microsoft On The Issues Asia Canada Europe, Middle East and Africa Latin America The Code of Us What's new AI Innovation Digital Transformation Sustainability Security Work & Life Diversity & Inclusion Unlocked Microsoft 365 Azure Copilot Windows Surface XBOX Deals Small Business Support Windows Apps Outlook OneDrive Microsoft Teams OneNote Microsoft Edge Moving from Skype to Teams Computers Shop XBOX Accessories VR & mixed reality Certified Refurbished Trade-in for cash XBOX Game Pass Ultimate PC Game Pass XBOX games PC games Microsoft AI Microsoft Security Dynamics 365 Microsoft 365 for business Microsoft Power Platform Windows 365 Small Business Digital Sovereignty Azure Microsoft Developer Microsoft Learn Support for AI marketplace apps Microsoft Tech Community Microsoft Marketplace Software companies Visual Studio Microsoft Rewards Free downloads & security Education Gift cards Licensing Unlocked stories View Sitemap

By builders, for builders.

A Microsoft publication

LLM-as-translator pattern: Adding natural language understanding to an existing system without rewriting it

Modernizing an existing implementation can be expensive and disruptive. Instead, consider introducing an abstraction layer to give your legacy system a modern facelift.

Many companies have legacy systems that still get the job done. But as technology advances, new innovations arise that those companies will want to leverage. 

Take conversational search. With the growing prevalence of LLMs, many consumers now expect the ability to have a natural-language conversation with search capabilities to help them find what they’re looking for, refining results along the way. However, most retailers already have search capabilities, often keyword-based, that work from existing catalogs. Often, these platforms support the rest of the business by providing APIs, which make them approachable and easy to use. 

The first instinct may be to replace those legacy systems, but modernizing the implementation can be expensive and disruptive to the business. 

One approach is to introduce an abstraction layer that enables a modern façade over an existing implementation. Rather than replacing the system, this layer translates the request into something that the system understands. 

While the use case discussed here is a search capability, this approach can be applied to any system that requires structured input.

Core principle

The system already knows how to answer a structured query. But it can’t parse a natural language request like “a cheap red kettle” into the required structured format:

query = "kettle" filters = color:red, price:0-30

Translation between a fuzzy human phrasing and a precise, structured format is exactly the kind of context-sensitive task LLMs are good at. It’s also a bounded task. The LLM never touches the data, never ranks results, and never decides what’s correct. It only produces input for a system that remains the source of truth. 

That boundary makes the pattern safe. If the translation is wrong, you get a slightly off query, not a corrupted system. And as I’ll show, you can detect and recover from a bad translation deterministically.

Architecture

The translator is a small service that sits between the caller and the existing implementation.

The LLM is never asked to invent structured terms. Instead, it’s given the actual vocabulary the API expects (e.g., valid filters). This eliminates most hallucinations. 

This layer can also downgrade to what the system did before, which means the fallback is baseline behavior.

Pipeline

A translation request flows through four stages.

Stage 0: The shape of the translations

Model the translated query as an explicit type rather than passing loose strings around.

public sealed record StructuredQuery { public required string Entity { get; init; } // the core noun, e.g. "kettle" public IReadOnlyList<Filter> Filters { get; init; } = []; // grounded, validated filters public string? Category { get; init; } // optional classification } public sealed record Filter(string Field, string Value);

Explicit types make validation, caching, observability, and versioning significantly easier than passing loosely structured prompts through the pipeline.

Stage 1: Normalize the request

Strip the modifiers to isolate the core intent and get the entity of interest. This is a lightweight task, so it can be routed to a small, inexpensive model. Keep the prompt narrow with strict output.

public async Task<string> ExtractEntityAsync(string rawQuery, CancellationToken ct) { var messages = new ChatMessage[] { new SystemChatMessage( "Identify the single product the user is searching for. " + "Return only that noun, with no extra words and no attributes " + "such as colour, size, or price."), new UserChatMessage(rawQuery), }; var response = await _cheapModel.CompleteChatAsync(messages, cancellationToken: ct); return response.Value.Content[0].Text.Trim(); }

“a cheap red kettle” ➡️ “kettle” 

Stage 2: Ground the model in real vocabulary

This stage is critical, and it’s one that might be missed. Don’t ask the LLM to guess which filters exist. Extract them from the original system and then pass this authoritative list to the model.  

It should be relatively simple to get the vocabulary from most systems. For example, search engines return facets, databases expose their schema, and APIs publish contracts. Once you have the vocabulary, prune it aggressively before sending it to the model. Remove values, nested children, and anything irrelevant. This reduces the number of tokens and can improve latency.

public async Task<string> GetGroundingVocabularyAsync(string entity, CancellationToken ct) { // Ask the engine itself what filters are valid for this entity. var facets = await _engine.GetAvailableFacetsAsync(entity, ct); var json = JObject.Parse(facets); // Prune high-volume, low-signal nodes to keep the prompt small and cheap. foreach (var path in new[] { "$..['values']", "$..['children']", "$..['id']" }) { foreach (var node in json.SelectTokens(path).ToArray()) node.Parent?.Remove(); } return json.ToString(Formatting.None); }

Stage 3: Translate with a strict, grounded prompt

This stage requires reasoning as it maps natural language onto a structured vocabulary. The prompt does three things: states the task, supplies the grounding data, and forbids intervention.

public async Task<IReadOnlyList<Filter>> TranslateFiltersAsync( string rawQuery, string vocabulary, CancellationToken ct) { var messages = new ChatMessage[] { new SystemChatMessage( "Translate the user's request into filters for a search engine. " + "Choose ONLY from the fields and values in the provided vocabulary; " + "do not invent fields or values that are not present. " + "[ SystemChatMessage the user's request"filters\":[{\"field\":\"\",\"value\":\"\"}]}. " + "Vocabulary: " + vocabulary), new UserChatMessage(rawQuery), }; var options = new ChatCompletionOptions { // Constrain the model to valid, parseable output. ResponseFormat = ChatResponseFormat.CreateJsonObjectFormat(), Temperature = 0f, }; var response = await _capableModel.CompleteChatAsync(messages, options, ct); var payload = JsonSerializer.Deserialize<FilterPayload>(response.Value.Content[0].Text); return payload?.Filters ?? []; } private sealed record FilterPayload(IReadOnlyList<Filter> Filters);

It’s worth highlighting the request for structured output and a Temperature setting of 0 to reduce creativity.  

The independent sub-translations have no dependency, so they can be run concurrently.

var filtersTask = TranslateFiltersAsync(rawQuery, vocabulary, ct); var categoryTask = ClassifyCategoryAsync(entity, ct); await Task.WhenAll(filtersTask, categoryTask); var query = new StructuredQuery { Entity = entity, Filters = await filtersTask, Category = await categoryTask, };

Stage 4: Execute with graceful degradation 

A successful translation could still be too narrow and return no results. In this use case, that wouldn’t be good. Therefore, begin with the full translated query and progressively relax it, ultimately ending with the untranslated query the system would have run anyway.

public async Task<SearchResult> ExecuteAsync( string rawQuery, StructuredQuery translated, CancellationToken ct) { // Ordered from most specific to the original baseline behaviour. var attempts = new Func<Task<SearchResult>>[] { () => _engine.SearchAsync(translated.Entity,translated.Category,translated.Filters,ct), () => _engine.SearchAsync(translated.Entity,category: null,translated.Filters,ct), () => _engine.SearchAsync(translated.Entity,translated.Category,filters: [],ct), () => _engine.SearchAsync(rawQuery,category: null,filters: [],ct), }; foreach (var attempt in attempts) { var result = await attempt(); if (result.HasHits) return result; } return SearchResult.Empty; }

This cascade makes the pattern operationally safe by providing a fallback to known behavior.

Caching the translation

Requests to the LLM are likely to incur cost and impact latency. In this case, this is the translation and not the search. Introducing a cache can deliver a better experience while reducing both cost and latency.  

The translated query should be cached, keyed by the normalized input, and then sent to the search as normal.

public async Task<StructuredQuery> TranslateAsync(string rawQuery, CancellationToken ct) { var key = Normalise(rawQuery); // lower-case, trim, collapse whitespace if (_cache.TryGetValue<StructuredQuery>(key, out var cached)) return cached!; var translated = await RunPipelineAsync(rawQuery, ct); _cache.Set(key, translated, new MemoryCacheEntryOptions { AbsoluteExpirationRelativeToNow = TimeSpan.FromHours(1), Size = 1, // bound the cache - never let it grow without limit }); return translated; }

Caching the translated query rather than the results ensures that the results are against the latest data.  

When to use this pattern  

Consider using this pattern when: 

Do not consider using this pattern when: 

The takeaway 

Beyond just making an existing system smarter, this pattern makes an existing system easier for people to use. By constraining the LLM to translation rather than decision making, organizations can introduce natural language experiences while retaining the reliability, governance, and business logic of the established systems they already trust.