Article

Beyond Context Engineering: The Case for Custom Models

Photo of David Meikle
David Meikle
|

September 29, 2026

Image via Unsplash+

At BloxWeaver, we build custom models for ourselves and our clients. We did not arrive at custom models because of the GenAI wave. We carried the belief into it.

In our previous businesses, we had already seen what happens when models are trained around real production work. Outputs get more accurate, edge cases get resolved, and quality starts to compound. Each correction has the potential to improve more than the item in front of the reviewer.

The lesson was clear - general models provide broad capability, but production precision comes from adapting the right model around the task, environment, and data. It is the foundation of what we call Precision AI.

That belief was not popular. I was told on CTO lists, at conferences, and in plenty of private conversations that training our own models was a waste of time. The advice was to wait for the next frontier release. The general model would catch up, and the custom work would become obsolete before it paid back.

Our experience has consistently been the opposite.

Frontier models have improved enormously, and we use them every day. At the same time, the custom models we build for ourselves and our clients are in production. Customised LLMs, translation, transcription, quality-estimation, and multimodal models work across language, audio, images, document layouts, and tightly bounded decisions. They deliver outcomes where generic models, longer prompts, and larger context windows did not meet the required standard.

A new frontier release raises the general baseline. It does not contain your reviewer corrections, domain terminology, language-pair characteristics, acoustic environment, visual categories, quality thresholds, or operating constraints, and it does not allow the learning from your production data to compound.

Waiting gives you a better general model. It does not make that model specific to your work.

Custom models are not simply a remedy for generic models that fall short. Where the task is understood, the data exists, and the performance or economics justify the investment, a custom model may be the right decision from the start.

The GenAI wave narrowed the model conversation

The arrival of ChatGPT changed what people thought was possible with AI. It also narrowed the conversation around how AI systems should be built. A much larger community began building AI-enabled solutions, often with a foundation-model API call as their main frame of reference.

Many projects started with the same assumption: take a general-purpose LLM, describe the task in a prompt, add examples, and keep expanding the context until the output is good enough.

Prompt engineering came first. Context engineering followed, bringing retrieval, memory, tool definitions, structured state, examples, and reference material into the context window. These techniques are real engineering disciplines, and we use both of them.

They are particularly effective when the model needs information that changes from one execution to the next. A customer record, a current product specification, a document being reviewed, or the result of a tool call belongs in the run-time context. The model needs that information now, not embedded permanently in its weights.

Problems appear when context becomes the answer to every limitation.

If a model repeatedly needs the same long instructions, corrections, examples, terminology rules, or decision patterns, the system is teaching it the same lesson on every call. The behaviour has not been learned. It is reconstructed each time, with the cost and variability that comes with that.

The context window is finite, and everything in it competes for attention. Every repeated instruction is paid for again on every call, and a provider update can quietly change behaviour a team spent months stabilising.

There is another problem hiding underneath. Sometimes the task should not belong to a general-purpose language model at all.

Not every AI problem is a language-generation problem

Production workflows rarely contain one task. They contain a collection of smaller tasks with different inputs, outputs, constraints, and definitions of success.

A multilingual video workflow, for example, may need to recognise speech, identify speakers, interpret what is visible on screen, translate dialogue, estimate translation quality, generate new speech, align it to timing constraints, and decide which items need human review.

A deterministic workflow may coordinate those steps, or an LLM may help where language, ambiguity, or judgement requires it. Neither choice means an LLM should perform every task.

  • Speech recognition models are designed to turn audio into text.
  • Machine translation models are trained on parallel text to map meaning from one language to another.
  • Quality-estimation models are designed to predict whether an output is likely to be good enough, without needing a reference to compare against.
  • Vision and layout models can reason over pixels, spatial relationships, pages, and interfaces.
  • Classifiers can make bounded decisions quickly and return confidence scores that can be calibrated into reliable probabilities, rather than generated prose.
  • Embedding models and rerankers can learn what relevance means within a particular domain.
  • LLMs are strong at language, judgement, ambiguity, tool use, and orchestration.

Choosing the right family matters because architectures, training objectives, and the training data shape what a model is good at. A larger prompt does not remove those differences.

A generic specialist model is still generic

A generic quality-estimation model that performs highly on a public machine translation benchmark does not automatically understand a client’s definition of quality. It may not reflect the language pair, subject matter, error categories, reviewer expectations, or risk threshold used in the real workflow.

The same applies across model types:

  • A generic speech model may struggle with specialist terminology, accents, recording conditions, or the way speakers use acronyms.
  • A generic translation model may not know the approved terminology, style, domain conventions, or language-pair characteristics that matter to the client.
  • A generic embedding model may capture broad semantic similarity without learning what relevance means in a specialist domain.

Vendors and labs make trade-offs for broad markets and product goals. Those choices may not match your task, and unless adaptation or controllability is part of the product, you have limited scope to change them.

The goal is not customisation for its own sake. The goal is to make the model fit the work.

Customisation is not new

The GenAI wave made model customisation sound like an alternative to the modern approach. In reality, adapting models around production environments was how much of production AI became useful in the first place.

Speech recognition systems were adapted to acoustic conditions, accents, and specialist vocabulary. Machine translation systems were tuned around language pairs, domains, terminology, and corrected translations. OCR and document models were trained around the layouts and document types they needed to process.

Leading AI labs make the same point through their own work. The strongest task-specific capabilities on public leaderboards do not come from prompting a base model alone. They come from post-training, including instruction tuning, preference optimisation, reinforcement learning, verifiers, and purpose-built training environments. Google DeepMind’s AlphaProof, for example, learned through reinforcement learning inside the Lean theorem prover, which checked each proof. The lesson is not that every production system needs this level of customisation. It is that serious performance gains come from adapting the model and its environment to the task.

The techniques continue to evolve, and the available base models are far more capable. The production principle remains familiar: start with general capability, then use representative data to close the gap between a generic model and the work it needs to perform.

What customisation actually means

The phrase “custom model” is often interpreted as a fine-tuned LLM. That is one option, although it is far from the only one.

Customisation can mean continued pretraining on domain or language data, fine-tuning on corrected examples, training a task-specific head (a small output layer trained on top of an existing model), learning from preferences, calibrating confidence scores, distilling capability from a larger model into a smaller specialist model (subject to the source model’s licence terms), or building a model around a tightly defined decision.

It can also mean combining these approaches. A translation system might use a customised MT model for the first pass, an LLM for automatic post-editing (APE) where richer context helps, a customised quality-estimation model to identify risk, and reviewer corrections to improve the next version of each model.

That distinction matters because precision is usually a property of the complete system, not one model in isolation.

Context and custom models often work together

Where a model can take context as input, context engineering remains essential. Context is the right place for information that is current, specific to this execution, or likely to change. Models are a better place for capabilities and behaviours that need to persist across executions.

However, context only helps if the model uses it well. A general-purpose model may ignore relevant information, give the wrong part too much weight, or apply it inconsistently across similar inputs.

We have often found that the best results come from combining context engineering with model customisation. The context supplies the information needed for the current task. Fine-tuning teaches the model which parts matter, how they relate to the task, and how they should influence the output.

Machine translation with automatic post-editing is a good example. An APE model may receive the source text, an initial translation, terminology, document context, quality signals, and previous corrections. Supplying those inputs is only part of the solution. The model also needs to learn when to preserve the initial translation, when to correct it, which contextual evidence should take priority, and how to avoid making unnecessary changes.

The same principle applies beyond translation. Models can be trained to use retrieved evidence more reliably, attend to relevant visual or document context, follow domain-specific metadata, interpret tool outputs, or apply contextual quality signals consistently.

Sometimes what belongs in the weights is the behaviour required to use the context properly.

A useful way to separate the responsibilities is:

Context

What the system needs now

Current facts, task state, retrieved knowledge, user inputs, examples, and reference material for this execution.

Tools

What the system can do

Deterministic operations, external actions, validation, data access, and specialist models exposed as controlled capabilities.

Models

What the system has learned

Persistent capabilities, domain patterns, decision boundaries, quality signals, and learned behaviour for interpreting context.

The boundaries are not fixed. Your data helps you decide where each behaviour belongs.

A new or changing instruction may belong in a prompt. Frequently updated knowledge belongs in retrieval, as fine-tuning is better at teaching behaviour than at storing facts reliably. A reliable external action belongs in a tool. A stable behaviour that appears repeatedly, can be measured, and has enough representative data may belong in a customised model.

This is where context engineering becomes part of the customisation story. A well-instrumented context-driven system generates examples, corrections, preferences, confidence signals, and outcomes. Those signals tell you which model to adapt and what it needs to learn.

When behaviour belongs in a model

Customisation does not need to follow months of prompting. When a task is stable, measurable, repeated, and supported by representative data, the behaviour may belong in a model from the start. The decision should still be supported by evidence.

We look for a combination of signals:

  • The task is stable. Its objective and success criteria are understood well enough to train against.
  • The result is measurable. You can evaluate whether the customised model is better than the baseline.
  • The work is repeated. The same capability or decision appears often enough for improvement to matter.
  • The data is representative. Corrections, labels, preferences, or outcomes reflect the production workload rather than an unrelated benchmark.
  • The available options leave a visible gap. Errors cluster around the domain, modality, language, terminology, or decision boundary you care about.

When those conditions are present, reconstructing the desired behaviour in every prompt may be the less practical design. The context gets longer, the instruction becomes more fragile, and every execution pays the cost again.

Keep the architecture as simple as the outcome allows. The right solution may be one specialist model integrated into existing software, a deterministic workflow combining several models, or an LLM or agent where language, ambiguity, and dynamic decisions genuinely require one. Add orchestration or agentic behaviour only when evaluation shows it improves the result, and if you do, we suggest you follow our Four Rules for Precision AI.

Performance comes first

Cost and sovereignty are valid reasons to explore custom models, but performance is where the decision should start.

The relevant comparison is not whether a custom model can beat a frontier model on a public leaderboard. It is whether the complete system performs better on the work you actually need to do.

That means measuring the things that matter in production: terminology accuracy, transcription errors, reviewer effort, decision precision and recall, quality-estimation calibration, turnaround time, exception rates, or whichever outcomes define success for the task.

Public benchmarks can help identify promising models, but they are narrow windows into performance: one dataset, one task definition, one scoring method. Your own evaluation set tells you whether a model truly works for you.

This is particularly important for multilingual models. An aggregate score can hide significant differences between languages and language pairs, especially where high-resource languages dominate the evaluation. Some multilingual benchmarks are also created through translation, which can flatten or distort the meaning they were intended to test. A model can perform well on the benchmark while still struggling with the terminology, linguistic variation, and quality expectations found in real work.

A model designed for a particular language pair, content type, quality policy, or production environment needs to be evaluated against that work.

Our approach is consistent even when the measurements differ. We establish a baseline using the strongest available models and techniques, and a customised model has to beat it on the same held-out work. If it does not, we use the stronger existing option.

Frontier progress should raise the baseline you test against, not erase the value of task-specific data.

Cost, control, and sovereignty follow

Once a customised model demonstrates an equivalent or better production outcome, the wider benefits become easier to evaluate.

A smaller model can be cheaper and faster than repeatedly sending a large prompt to a frontier model, and reasoning models widen that gap because the same visible output can carry many times the billed tokens. A model deployed in your own environment can give you more control over data handling, availability, versioning, and upstream policy changes. A system with clearly defined responsibilities can also make failures easier to diagnose than one model attempting the whole workflow in a single generation.

These benefits depend on the workload. Training, hosting, monitoring, and maintaining models all have costs. The answer should come from comparing the complete production system, not token prices or training costs in isolation.

Volume typically changes that calculation dramatically. A customisation effort that makes little sense for a few hundred calls may become compelling when the same task runs millions of times and every improvement compounds.

A practical model strategy

Using custom models does not require starting with a large training programme. The sensible route begins with the task and the assets already available:

A practical model strategy

1. Decompose the work

Understand the real tasks

Identify the inputs, outputs, modalities, decisions, failure modes, and quality requirements within the workflow.

2. Assess the starting position

Understand the assets and constraints

Review the available data, expected volume, budget, latency, sovereignty requirements, and how long the capability will be used.

3. Choose the model strategy

Select for task fit

Consider generic and custom options from the outset, then decide which model families or deterministic software should own the work.

4. Define the comparison

Set the baseline and evaluation

Compare the viable approaches on representative data using the quality, cost, latency, and control measures that matter to the task.

5. Build the right solution

Train from the outset where appropriate

Use the strongest generic option where it fits, or customise the appropriate model where the task, data, and economics support it.

6. Instrument and re-evaluate

Keep proving the choice

Capture production signals, compare new options, recalibrate decisions, and retain each model only while it continues to improve the outcome.

Running the system should generate the evidence and data needed to improve its next version.

Different signals support different improvements. Corrections create supervised examples, rankings support preference tuning, verifiable outcomes can become rewards, and reviewer decisions improve quality-estimation and calibration models. Clusters of similar failures may show that the system needs a different model family rather than another instruction in the prompt.

More practical than it used to be

The reasonable objection is that custom model development sounds like work reserved for frontier labs and very large engineering teams. The barriers are lower than that assumption suggests.

  • Data need not be the blocker. Existing operational data can provide training and evaluation signals. Where coverage is limited, carefully generated and validated synthetic data can add representative examples for edge cases, underrepresented languages, specialist domains, rare conditions, or difficult modalities. Production then extends that evidence over time.
  • Adaptation does not always require training from scratch. Open-weight models, adapters, continued pretraining, fine-tuning services, and specialist frameworks provide several starting points. Adapters (small trainable layers, such as LoRA, added to a frozen base model) can often be retrained on a newer base model using the same data.
  • Deployment is a spectrum. A model can run through a managed endpoint, dedicated infrastructure, rented hardware, or an organisation’s own environment depending on its scale and sovereignty requirements.
  • Evaluation limits the risk. A representative evaluation set lets a team compare the customised model with its existing baseline before changing the production workflow.

The work still needs a justified outcome, but teams can now test the case incrementally rather than treating custom model development as one large, irreversible investment.

What this means for production teams

Context engineering remains powerful, and it generates the signals that tell you where to customise. The mistake is treating it as the destination for every behaviour and every model type.

Where good enough is genuinely good enough, a generic model with a better prompt may be the right answer. Where accuracy matters, a custom model from day one can be your best option, whether an adapted speech or translation model, a custom quality-estimation model, a vision classifier, a domain-specific embedding model, or a fine-tuned LLM.

The advantage comes from knowing which part of the system should own which behaviour, and improving each part against the outcome that matters.

Ease of access is not the same as fitness for production.

If you would like to discuss where custom models fit your AI strategy and solutions, what data you already have, or how to prove the performance case before committing to training, our model training team can help.

To stay up to date on our latest blog posts, you can subscribe to our mailing list and/or follow us on LinkedIn.

Want to connect?

Book a meeting with BloxWeaver to find time to talk through production AI, localization workflows, and multilingual content operations.