Article
September 29, 2026
At BloxWeaver, we build custom models for ourselves and our clients. We did not arrive at custom models because of the GenAI wave. We carried the belief into it.
In our previous businesses, we had already seen what happens when models are trained around real production work. Outputs get more accurate, edge cases get resolved, and quality starts to compound. Each correction has the potential to improve more than the item in front of the reviewer.
The lesson was clear - general models provide broad capability, but production precision comes from adapting the right model around the task, environment, and data. It is the foundation of what we call Precision AI.
That belief was not popular. I was told on CTO lists, at conferences, and in plenty of private conversations that training our own models was a waste of time. The advice was to wait for the next frontier release. The general model would catch up, and the custom work would become obsolete before it paid back.
Our experience has consistently been the opposite.
Frontier models have improved enormously, and we use them every day. At the same time, the custom models we build for ourselves and our clients are in production. Customised LLMs, translation, transcription, quality-estimation, and multimodal models work across language, audio, images, document layouts, and tightly bounded decisions. They deliver outcomes where generic models, longer prompts, and larger context windows did not meet the required standard.
A new frontier release raises the general baseline. It does not contain your reviewer corrections, domain terminology, language-pair characteristics, acoustic environment, visual categories, quality thresholds, or operating constraints, and it does not allow the learning from your production data to compound.
Waiting gives you a better general model. It does not make that model specific to your work.
Custom models are not simply a remedy for generic models that fall short. Where the task is understood, the data exists, and the performance or economics justify the investment, a custom model may be the right decision from the start.
The arrival of ChatGPT changed what people thought was possible with AI. It also narrowed the conversation around how AI systems should be built. A much larger community began building AI-enabled solutions, often with a foundation-model API call as their main frame of reference.
Many projects started with the same assumption: take a general-purpose LLM, describe the task in a prompt, add examples, and keep expanding the context until the output is good enough.
Prompt engineering came first. Context engineering followed, bringing retrieval, memory, tool definitions, structured state, examples, and reference material into the context window. These techniques are real engineering disciplines, and we use both of them.
They are particularly effective when the model needs information that changes from one execution to the next. A customer record, a current product specification, a document being reviewed, or the result of a tool call belongs in the run-time context. The model needs that information now, not embedded permanently in its weights.
Problems appear when context becomes the answer to every limitation.
If a model repeatedly needs the same long instructions, corrections, examples, terminology rules, or decision patterns, the system is teaching it the same lesson on every call. The behaviour has not been learned. It is reconstructed each time, with the cost and variability that comes with that.
The context window is finite, and everything in it competes for attention. Every repeated instruction is paid for again on every call, and a provider update can quietly change behaviour a team spent months stabilising.
There is another problem hiding underneath. Sometimes the task should not belong to a general-purpose language model at all.
Production workflows rarely contain one task. They contain a collection of smaller tasks with different inputs, outputs, constraints, and definitions of success.
A multilingual video workflow, for example, may need to recognise speech, identify speakers, interpret what is visible on screen, translate dialogue, estimate translation quality, generate new speech, align it to timing constraints, and decide which items need human review.
A deterministic workflow may coordinate those steps, or an LLM may help where language, ambiguity, or judgement requires it. Neither choice means an LLM should perform every task.
Choosing the right family matters because architectures, training objectives, and the training data shape what a model is good at. A larger prompt does not remove those differences.
A generic quality-estimation model that performs highly on a public machine translation benchmark does not automatically understand a client’s definition of quality. It may not reflect the language pair, subject matter, error categories, reviewer expectations, or risk threshold used in the real workflow.
The same applies across model types:
Vendors and labs make trade-offs for broad markets and product goals. Those choices may not match your task, and unless adaptation or controllability is part of the product, you have limited scope to change them.
The goal is not customisation for its own sake. The goal is to make the model fit the work.
The GenAI wave made model customisation sound like an alternative to the modern approach. In reality, adapting models around production environments was how much of production AI became useful in the first place.
Speech recognition systems were adapted to acoustic conditions, accents, and specialist vocabulary. Machine translation systems were tuned around language pairs, domains, terminology, and corrected translations. OCR and document models were trained around the layouts and document types they needed to process.
Leading AI labs make the same point through their own work. The strongest task-specific capabilities on public leaderboards do not come from prompting a base model alone. They come from post-training, including instruction tuning, preference optimisation, reinforcement learning, verifiers, and purpose-built training environments. Google DeepMind’s AlphaProof, for example, learned through reinforcement learning inside the Lean theorem prover, which checked each proof. The lesson is not that every production system needs this level of customisation. It is that serious performance gains come from adapting the model and its environment to the task.
The techniques continue to evolve, and the available base models are far more capable. The production principle remains familiar: start with general capability, then use representative data to close the gap between a generic model and the work it needs to perform.
The phrase “custom model” is often interpreted as a fine-tuned LLM. That is one option, although it is far from the only one.
Customisation can mean continued pretraining on domain or language data, fine-tuning on corrected examples, training a task-specific head (a small output layer trained on top of an existing model), learning from preferences, calibrating confidence scores, distilling capability from a larger model into a smaller specialist model (subject to the source model’s licence terms), or building a model around a tightly defined decision.
It can also mean combining these approaches. A translation system might use a customised MT model for the first pass, an LLM for automatic post-editing (APE) where richer context helps, a customised quality-estimation model to identify risk, and reviewer corrections to improve the next version of each model.
That distinction matters because precision is usually a property of the complete system, not one model in isolation.
Where a model can take context as input, context engineering remains essential. Context is the right place for information that is current, specific to this execution, or likely to change. Models are a better place for capabilities and behaviours that need to persist across executions.
However, context only helps if the model uses it well. A general-purpose model may ignore relevant information, give the wrong part too much weight, or apply it inconsistently across similar inputs.
We have often found that the best results come from combining context engineering with model customisation. The context supplies the information needed for the current task. Fine-tuning teaches the model which parts matter, how they relate to the task, and how they should influence the output.
Machine translation with automatic post-editing is a good example. An APE model may receive the source text, an initial translation, terminology, document context, quality signals, and previous corrections. Supplying those inputs is only part of the solution. The model also needs to learn when to preserve the initial translation, when to correct it, which contextual evidence should take priority, and how to avoid making unnecessary changes.
The same principle applies beyond translation. Models can be trained to use retrieved evidence more reliably, attend to relevant visual or document context, follow domain-specific metadata, interpret tool outputs, or apply contextual quality signals consistently.
Sometimes what belongs in the weights is the behaviour required to use the context properly.
A useful way to separate the responsibilities is:
Context
Current facts, task state, retrieved knowledge, user inputs, examples, and reference material for this execution.
Tools
Deterministic operations, external actions, validation, data access, and specialist models exposed as controlled capabilities.
Models
Persistent capabilities, domain patterns, decision boundaries, quality signals, and learned behaviour for interpreting context.
The boundaries are not fixed. Your data helps you decide where each behaviour belongs.
A new or changing instruction may belong in a prompt. Frequently updated knowledge belongs in retrieval, as fine-tuning is better at teaching behaviour than at storing facts reliably. A reliable external action belongs in a tool. A stable behaviour that appears repeatedly, can be measured, and has enough representative data may belong in a customised model.
This is where context engineering becomes part of the customisation story. A well-instrumented context-driven system generates examples, corrections, preferences, confidence signals, and outcomes. Those signals tell you which model to adapt and what it needs to learn.
Customisation does not need to follow months of prompting. When a task is stable, measurable, repeated, and supported by representative data, the behaviour may belong in a model from the start. The decision should still be supported by evidence.
We look for a combination of signals:
When those conditions are present, reconstructing the desired behaviour in every prompt may be the less practical design. The context gets longer, the instruction becomes more fragile, and every execution pays the cost again.
Keep the architecture as simple as the outcome allows. The right solution may be one specialist model integrated into existing software, a deterministic workflow combining several models, or an LLM or agent where language, ambiguity, and dynamic decisions genuinely require one. Add orchestration or agentic behaviour only when evaluation shows it improves the result, and if you do, we suggest you follow our Four Rules for Precision AI.
Cost and sovereignty are valid reasons to explore custom models, but performance is where the decision should start.
The relevant comparison is not whether a custom model can beat a frontier model on a public leaderboard. It is whether the complete system performs better on the work you actually need to do.
That means measuring the things that matter in production: terminology accuracy, transcription errors, reviewer effort, decision precision and recall, quality-estimation calibration, turnaround time, exception rates, or whichever outcomes define success for the task.
Public benchmarks can help identify promising models, but they are narrow windows into performance: one dataset, one task definition, one scoring method. Your own evaluation set tells you whether a model truly works for you.
This is particularly important for multilingual models. An aggregate score can hide significant differences between languages and language pairs, especially where high-resource languages dominate the evaluation. Some multilingual benchmarks are also created through translation, which can flatten or distort the meaning they were intended to test. A model can perform well on the benchmark while still struggling with the terminology, linguistic variation, and quality expectations found in real work.
A model designed for a particular language pair, content type, quality policy, or production environment needs to be evaluated against that work.
Our approach is consistent even when the measurements differ. We establish a baseline using the strongest available models and techniques, and a customised model has to beat it on the same held-out work. If it does not, we use the stronger existing option.
Frontier progress should raise the baseline you test against, not erase the value of task-specific data.
Once a customised model demonstrates an equivalent or better production outcome, the wider benefits become easier to evaluate.
A smaller model can be cheaper and faster than repeatedly sending a large prompt to a frontier model, and reasoning models widen that gap because the same visible output can carry many times the billed tokens. A model deployed in your own environment can give you more control over data handling, availability, versioning, and upstream policy changes. A system with clearly defined responsibilities can also make failures easier to diagnose than one model attempting the whole workflow in a single generation.
These benefits depend on the workload. Training, hosting, monitoring, and maintaining models all have costs. The answer should come from comparing the complete production system, not token prices or training costs in isolation.
Volume typically changes that calculation dramatically. A customisation effort that makes little sense for a few hundred calls may become compelling when the same task runs millions of times and every improvement compounds.
Using custom models does not require starting with a large training programme. The sensible route begins with the task and the assets already available:
A practical model strategy
1. Decompose the work
Identify the inputs, outputs, modalities, decisions, failure modes, and quality requirements within the workflow.
2. Assess the starting position
Review the available data, expected volume, budget, latency, sovereignty requirements, and how long the capability will be used.
3. Choose the model strategy
Consider generic and custom options from the outset, then decide which model families or deterministic software should own the work.
4. Define the comparison
Compare the viable approaches on representative data using the quality, cost, latency, and control measures that matter to the task.
5. Build the right solution
Use the strongest generic option where it fits, or customise the appropriate model where the task, data, and economics support it.
6. Instrument and re-evaluate
Capture production signals, compare new options, recalibrate decisions, and retain each model only while it continues to improve the outcome.
Running the system should generate the evidence and data needed to improve its next version.
Different signals support different improvements. Corrections create supervised examples, rankings support preference tuning, verifiable outcomes can become rewards, and reviewer decisions improve quality-estimation and calibration models. Clusters of similar failures may show that the system needs a different model family rather than another instruction in the prompt.
The reasonable objection is that custom model development sounds like work reserved for frontier labs and very large engineering teams. The barriers are lower than that assumption suggests.
The work still needs a justified outcome, but teams can now test the case incrementally rather than treating custom model development as one large, irreversible investment.
Context engineering remains powerful, and it generates the signals that tell you where to customise. The mistake is treating it as the destination for every behaviour and every model type.
Where good enough is genuinely good enough, a generic model with a better prompt may be the right answer. Where accuracy matters, a custom model from day one can be your best option, whether an adapted speech or translation model, a custom quality-estimation model, a vision classifier, a domain-specific embedding model, or a fine-tuned LLM.
The advantage comes from knowing which part of the system should own which behaviour, and improving each part against the outcome that matters.
Ease of access is not the same as fitness for production.
If you would like to discuss where custom models fit your AI strategy and solutions, what data you already have, or how to prove the performance case before committing to training, our model training team can help.
To stay up to date on our latest blog posts, you can subscribe to our mailing list and/or follow us on LinkedIn.
Book a meeting with BloxWeaver to find time to talk through production AI, localization workflows, and multilingual content operations.