NearSync Help

NearSync Search

Speaking to It Dictating into the box, and how that differs from a live spoken conversation.

There are two quite different ways of using your voice here, and confusing them is the main reason people find it inconsistent.

Dictation: your voice becomes text

The first is straightforward. You speak, and what you said appears in the box as if you had typed it.

Hold ⌘K for a moment and dictation starts. Release, and what you said is transcribed into the search box.

The microphone button inside the box does the same thing with a click rather than a hold.

Hold ⌘; to dictate into the capture lane instead, for creating rather than searching.

The hold is deliberate rather than a quirk. A brief press of ⌘K opens the box, and holding it for a moment longer starts listening, so one chord does both without a separate shortcut to remember.

What happens to what you said

It lands as ordinary text and is treated as ordinary text.

That matters. A dictated sentence classifies exactly like a typed one: a question offers to ask, something with a time offers to create, a name searches. Speaking does not force your words into a task, and it does not send anything to the assistant.

If there is already something in the box, what you say is appended rather than replacing it, so you can type a fragment and finish it out loud.

Live voice: a spoken conversation

The second is different in kind. This is a conversation, out loud, in both directions.

Where your organisation has it enabled, the control in the search bar at the top of the screen opens the box straight into a live session. The assistant listens, answers aloud, reads your live data mid-answer, and can make changes if you agree to them out loud.

The full behaviour is covered in the Workspace piece on talking out loud.

Telling them apart

Dictation Live voice
What it does Turns speech into text Holds a conversation
Where it lands The box, as typing The panel, as an exchange
Does it answer No Yes, aloud
Cost Small, one transcription Higher, charged on audio both ways
Availability Always Only if enabled

Dictation is an input method. Live voice is a mode of working.

If what you want is to get words into the box without typing, use dictation. It is cheaper, faster, and it does not commit you to anything.

When dictation is genuinely better

Long capture. A sentence with a time, two people and a reminder is slow to type and quick to say.

On a phone. Where typing is worst, this is at its best.

When your hands are busy but the screen is not, which is a common enough state at a desk with a document open.

When it is worse

Anything with a name in it that is not common. Transcription is confident and wrong about unusual names, and you will spend more time correcting than typing.

Noisy rooms. Accuracy falls quickly, and a garbled capture is worse than none because you have to work out what you meant.

Short queries. Saying a three-letter search term is slower than typing it.

Accuracy, honestly

Transcription is good at ordinary sentences and predictable at what it gets wrong.

Common words are reliable. A sentence of everyday English comes back accurately most of the time.

Names are not. Anything outside a common English name list is transcribed phonetically, and confidently. Colleagues' names improve with use; customers' names frequently do not.

Days are the classic error. Tuesday and Thursday, in particular, and the consequence lands in a calendar rather than in a sentence.

Product jargon is the weakest. Module names, abbreviations and internal shorthand appear nowhere in general language and come back as something else entirely.

The mitigation is the same in every case: the text lands in the box, visible, before anything happens with it. Read it. Correcting one word is faster than undoing a task with the wrong date.

Privacy, briefly

Speaking your organisation's business out loud is a decision about the room you are in rather than about the software. Worth one thought before dictating a customer name in an open office or on a train.

The microphone is never open without a visible indicator, and closing the box stops any recording immediately.

Which chord for which job

The two hold shortcuts do genuinely different things and the difference is worth internalising.

Hold ⌘K puts your words in the search box. They are then classified exactly as typing would be: a name searches, a question offers to ask, a sentence with a time offers to create. Nothing is decided for you.

Hold ⌘; puts your words straight into the capture composer, where the at-verbs and the date parsing are already engaged. Use this when you know you are creating something.

Choosing wrongly is harmless. Words dictated into search that turn out to be a task are one row away from becoming one.

When it does not start

Two causes, and they look identical from the outside.

Microphone permission. The browser asks once and remembers. If it was refused, nothing happens when you hold the chord, and it has to be granted from the browser's own site settings rather than from here.

No microphone available. A machine with no input device, or one already held exclusively by another application.

In both cases the box carries on working normally for typing, which is why the failure is quiet rather than an error.

Practical notes

The microphone never stays open. Closing the box stops any recording in progress, and there is no state in which it is listening without a visible indicator.

Transcription takes a moment. A spinner replaces the microphone while it runs, and the text lands when it finishes.

Check before you commit. Dictated text sits in the box like anything else, so read it before pressing Enter, particularly when a date is involved.

If nothing happens, the browser has not been given microphone access. That is a permission prompt rather than a fault, and it is granted once per browser.

5 minUpdated 28 July 2026

Did this answer your question?

No, ask a person