When I tested Qwen3.8-27B on my original M1 Max Mac Studio in August, the number that changed my mind was not generation speed. It was the wait before the first answer appeared. The model could write at around 10–12 tokens a second, but a 55k-token prompt left me waiting more than 12 minutes before it started.
I have since loaded qwen3.6:35b-a3b-q4_K_M in Ollama on the same Mac Studio. It is a different model, using a mixture-of-experts architecture rather than the dense 27B model I tested before. I wanted to know whether it was actually more pleasant to use on this machine, not just whether it produced a larger tokens-per-second number.
So I ran the tests on the Mac Studio itself. These are measurements from my machine, not figures copied from a model card.
My setup and how I timed it
The machine is a first-generation Mac Studio M1 Max with 64 GB of unified memory. It was running Ollama 0.34.4 with the Q4_K_M model loaded at a 65,536-token context. Ollama reported about 24 GB resident and 100% GPU for the model. The model was already loaded before timing began, so these numbers do not include a cold start.
I used Ollama's streaming chat API with temperature 0. For the speed and retrieval runs I set
Short prompts: around 63 tokens a second
I gave it a short canonical-tag explanation, a Python URL-grouping task and an ecommerce SEO investigation prompt. These were real response prompts, not just “write the word hello” repeated until I liked the number.
That is a much better interactive feel than I got from the dense model in August. Once it starts answering, the text arrives faster than I can comfortably read it. The six generation measurements were also tightly grouped, from 62.43 to 63.42 tok/s.
There is an important qualification, though: speed is not quality. I did not execute unit tests against the Python answer, so I am not awarding it a coding pass. And the SEO answer looked plausible while recommending Search Console's old URL Parameters tool, which Google retired in 2022. That is exactly the sort of confident, out-of-date advice I would not want to apply to a client's site without checking.
The longer the prompt, the more the wait matters
I then used a synthetic audit file: hundreds of ordinary product-page records, with one randomly generated
It found all nine codes, including placements near the beginning, middle and end of the 54.6k-token prompt. I saw no retrieval failure in these runs. I also saw no increase in the Mac's reported swap use while these tests ran, although the machine already had swap in use before I started.
But 54.6k tokens still means sitting there for roughly two minutes and 20 seconds before the answer appears. That is a considerable improvement in my experience, yet it is not what I would call instant. The same lesson from the earlier benchmark survives: a model's writing speed is not the same as the speed at which a conversation feels responsive.
A reasoning limit that looked like a model failure
I also tried a small, verifiable URL-counting question at low, medium and xhigh reasoning settings. There are 85 valid URLs in the question and 15 duplicates, so the answer is 70. With a 1,024-token output allowance, low and xhigh used the whole budget on thinking and returned no visible answer. Medium began the correct answer but was cut off. All three requests ended because they hit the length limit.
That initially looks like a failure if you score only the final text. I reran medium with a 4,096-token allowance. It stopped normally after 20.4 seconds and returned the full, correct answer: 70 distinct valid URLs. This is almost a repeat of something I found in my earlier Qwen3.8 benchmark: a benchmark harness can make a reasoning model look broken simply by not leaving it enough room to finish.
What I would actually use it for
On this Mac Studio, I would happily use this Ollama setup for short, local drafting and exploratory work. It feels responsive in a way my previous test did not. I can also see a place for longer, unattended jobs: pull a specific item out of a large audit export, process documentation, or help triage a crawl while I do something else.
I would not treat the 9/9 retrieval score as proof that it understands a messy, real-world 55k-token project. My filler records were repetitive, the answer was an exact string, and most cells ran once. Nor would I let it make SEO implementation decisions unsupervised; the obsolete Search Console recommendation is a good reminder of why.
It is tempting to put “63 tok/s versus 10–12 tok/s” or “2 minutes versus 13 minutes” in a headline and call one model six times faster. I don't think that would be fair. The models have different architectures, the Ollama versions and settings changed, and the long prompts were not identical. These are two sets of observations on the same Mac Studio, not a controlled model-vs-model experiment.
My answer to the question I started with is simpler: this is the first Qwen setup I have run locally on the M1 Max that I would genuinely choose for a quick interaction. The big-context use case is much better, but I would still run it as a background job rather than sit and watch the cursor. If anyone else is running this tag on an original Mac Studio, I would be interested in your prompt-processing times as much as your generation speeds.
I have since loaded qwen3.6:35b-a3b-q4_K_M in Ollama on the same Mac Studio. It is a different model, using a mixture-of-experts architecture rather than the dense 27B model I tested before. I wanted to know whether it was actually more pleasant to use on this machine, not just whether it produced a larger tokens-per-second number.
So I ran the tests on the Mac Studio itself. These are measurements from my machine, not figures copied from a model card.
My setup and how I timed it
The machine is a first-generation Mac Studio M1 Max with 64 GB of unified memory. It was running Ollama 0.34.4 with the Q4_K_M model loaded at a 65,536-token context. Ollama reported about 24 GB resident and 100% GPU for the model. The model was already loaded before timing began, so these numbers do not include a cold start.
I used Ollama's streaming chat API with temperature 0. For the speed and retrieval runs I set
think: false; reasoning gets its own test below. I recorded the time until the first visible answer, along with Ollama's prompt-evaluation and generation durations. All measured requests reported zero cached prompt tokens. For the short prompts I ran each task twice. Most long-context cells were one run, so please don't read decimal-place precision as a promise of repeatability.Short prompts: around 63 tokens a second
I gave it a short canonical-tag explanation, a Python URL-grouping task and an ecommerce SEO investigation prompt. These were real response prompts, not just “write the word hello” repeated until I liked the number.
| Task | Run 1 | Run 2 | Time to first content |
|---|---|---|---|
| Short explanation | 63.16 tok/s | 63.15 tok/s | 0.53s / 0.26s |
| Python function | 63.42 tok/s | 63.37 tok/s | 0.29s / 0.29s |
| SEO investigation | 62.95 tok/s | 62.43 tok/s | 0.31s / 0.31s |
There is an important qualification, though: speed is not quality. I did not execute unit tests against the Python answer, so I am not awarding it a coding pass. And the SEO answer looked plausible while recommending Search Console's old URL Parameters tool, which Google retired in 2022. That is exactly the sort of confident, out-of-date advice I would not want to apply to a client's site without checking.
The longer the prompt, the more the wait matters
I then used a synthetic audit file: hundreds of ordinary product-page records, with one randomly generated
OVERRIDE_CODE hidden at different depths. Qwen had to return only that exact code. It is a deliberately simple retrieval test, but it lets me check both whether the model can find the information and how long it takes before I see it.| Prompt size | Retrieval | Prompt processing | First content |
|---|---|---|---|
| ~3,200 tokens | 1/1 | 4.6s | 4.6s |
| ~9,400 tokens | 1/1 | 14.5s | 14.5s |
| ~15,600 tokens | 3/3 | ~26.1s | ~26.2s |
| ~31,200 tokens | 1/1 | 63.1s | 63.3s |
| ~54,600 tokens | 3/3 | ~139.7s | 140.0–143.9s |
But 54.6k tokens still means sitting there for roughly two minutes and 20 seconds before the answer appears. That is a considerable improvement in my experience, yet it is not what I would call instant. The same lesson from the earlier benchmark survives: a model's writing speed is not the same as the speed at which a conversation feels responsive.
A reasoning limit that looked like a model failure
I also tried a small, verifiable URL-counting question at low, medium and xhigh reasoning settings. There are 85 valid URLs in the question and 15 duplicates, so the answer is 70. With a 1,024-token output allowance, low and xhigh used the whole budget on thinking and returned no visible answer. Medium began the correct answer but was cut off. All three requests ended because they hit the length limit.
That initially looks like a failure if you score only the final text. I reran medium with a 4,096-token allowance. It stopped normally after 20.4 seconds and returned the full, correct answer: 70 distinct valid URLs. This is almost a repeat of something I found in my earlier Qwen3.8 benchmark: a benchmark harness can make a reasoning model look broken simply by not leaving it enough room to finish.
What I would actually use it for
On this Mac Studio, I would happily use this Ollama setup for short, local drafting and exploratory work. It feels responsive in a way my previous test did not. I can also see a place for longer, unattended jobs: pull a specific item out of a large audit export, process documentation, or help triage a crawl while I do something else.
I would not treat the 9/9 retrieval score as proof that it understands a messy, real-world 55k-token project. My filler records were repetitive, the answer was an exact string, and most cells ran once. Nor would I let it make SEO implementation decisions unsupervised; the obsolete Search Console recommendation is a good reminder of why.
It is tempting to put “63 tok/s versus 10–12 tok/s” or “2 minutes versus 13 minutes” in a headline and call one model six times faster. I don't think that would be fair. The models have different architectures, the Ollama versions and settings changed, and the long prompts were not identical. These are two sets of observations on the same Mac Studio, not a controlled model-vs-model experiment.
My answer to the question I started with is simpler: this is the first Qwen setup I have run locally on the M1 Max that I would genuinely choose for a quick interaction. The big-context use case is much better, but I would still run it as a background job rather than sit and watch the cursor. If anyone else is running this tag on an original Mac Studio, I would be interested in your prompt-processing times as much as your generation speeds.