NearSync Help

AI

Choosing Your Model Which model answers, why a fallback matters more than the primary, and what temperature actually changes.

Model settings is where you decide which model answers. Providers made capability available; this chooses from it.

The choice that matters

Models differ in three ways that show up in daily use.

How good they are at reasoning. Noticeable on anything involving several steps or a judgement.

How fast they answer. Noticeable on everything, and more than people expect. A model that is slightly better and twice as slow is usually the wrong choice for an assistant somebody uses forty times a day.

What they cost to run, in credits or in money depending on your setup.

There is no model that wins all three, which is the entire reason this page exists.

Pick for the common case, not the hard one

The usual mistake is choosing the most capable model available because it handles the hardest question best.

Most questions are not hard. They are "what is this deal worth", "when did we last speak to them", "summarise this thread". A fast model handles those perfectly, and the assistant feels alive rather than thoughtful.

Choose for the ninetieth question, not the first. The hard one can go to a better model deliberately.

Set a fallback

The most valuable setting on the page, and the one people skip.

If the primary model is unavailable, the fallback answers. Without one, the assistant simply stops, and it stops in the least helpful way: silently, mid-conversation, for reasons nobody in the room can diagnose.

Put the fallback on a different provider from the primary. A fallback that shares a vendor with the primary fails at exactly the same moment as the primary. That is not a fallback, it is a second copy of the same risk.

Temperature

One control, badly named, and worth understanding because people set it wrong in both directions.

Low produces consistent, predictable answers. The same question twice gets nearly the same answer.

High produces varied, more creative answers. The same question twice gets two different ones.

Low is right for almost everything here. This is a business assistant answering questions about your data. Variety is not a virtue when somebody asks what an invoice is worth.

Raise it only for genuinely generative work, like drafting several versions of marketing copy, and only where the person asking wants options rather than an answer.

Different models for different work

Where it is supported, matching the model to the job is worth doing.

Answering questions about data wants fast and accurate, not imaginative.

Drafting and rewriting benefits from a stronger model, because the output is read by a customer.

Classifying and routing wants the cheapest model that is reliable, because it happens constantly and nobody reads the intermediate result.

The instinct to use the best model everywhere is expensive and, on the third of those, actively pointless.

Voice

Speaking to the assistant uses a purpose-built voice model rather than the text one, because a conversation out loud has different requirements: it has to respond immediately, handle interruption, and sound like something a person can talk over.

It is configured separately. If voice does not work but typing does, that is where to look rather than at the text model. Using it is covered in Talking out loud and Speaking to search.

Context length

How much the model can hold at once, and the reason a long conversation sometimes seems to forget its own beginning.

It matters for long documents and long threads. Summarising a forty-page contract or a thread with sixty messages needs a model that can take it all in; a smaller one silently works from a portion and produces a confident summary of the part it saw.

It matters much less for ordinary questions, which is most of them.

If people report that the assistant loses the plot on long material and is fine otherwise, this is usually why, and it is a model choice rather than anything wrong.

Streaming

Answers appear as they are produced rather than arriving complete.

That is worth knowing because it changes how latency feels. A model that takes six seconds to finish but starts in one feels responsive; a model that takes four seconds and shows nothing until the end feels slow. The measured numbers will disagree with what people report, and people are describing something real.

When judging a model, watch when the first words appear, not only when the last do.

Before you change anything

Test the new model against the old one on questions you actually ask. Comparing models runs both on the same prompt side by side, which takes about a minute and settles the argument with evidence rather than impression.

Change one thing. Switching the model and raising the temperature together produces a difference nobody can attribute.

Then watch the error rate and latency in usage for a few days. A model that answers beautifully and takes nine seconds will be quietly abandoned by everybody, and the usage numbers are where you find that out.

Who this affects

One setting, everybody.

The model chosen here answers for the whole organisation, in every surface: the cockpit, the command bar, chat, drafted replies, and anything an agent writes.

That is deliberate. Per-person models would mean two colleagues asking the same question getting different answers of different quality, which is the fastest way to make an assistant untrustworthy.

It also means a change is felt everywhere at once, which is the argument for testing before switching rather than after.

A note on new releases

A newer model is not automatically better for your work. It is better on the benchmarks its maker published.

Test it the same way you would test any other change, on your own questions, before it becomes the one everybody depends on.

What this page will not fix

Answers that are wrong about your organisation. That is a knowledge problem, not a model problem, and it is solved in Teaching the assistant. A better model with no knowledge of your business gives a more articulate wrong answer.

Answers in the wrong tone. That is Persona and guardrails.

5 minUpdated 28 July 2026

Did this answer your question?

No, ask a person