AIに実務コードを頼んだら、正解でも次の回には答えが返らなかった
AIを仕事に使うとき、どのモデルを選ぶか。
Senteで使うモデルを選ぶために、短い実務課題を作って比べています。今回はDeepSeekとGPTの優劣を最初から決めず、実際に結果が分かれる問題を探しました。
使ったAPIのモデル名は DeepSeek V4.1 Flash(deepseek/deepseek-v4.1-flash)と GPT-6 Astra(openai/gpt-6-astra)です。
差が見つかったのは、答えが返ってくるかどうかでした。同じ問題で、正しく動くコードが返る回と、本文が空になる回もありました。
1つ目は、小数の丸め
作ってもらったのは、数字の文字列を小数第2位に丸めるJavaScript関数です。
条件は「ちょうど中間なら、残す最下位の桁を偶数にする」。たとえば、こうなります。
| 入力 | 正解 |
1.005 | 1.00 |
1.015 | 1.02 |
-1.015 | -1.02 |
-0.005 | 0.00 |
999999999999999999.995 | 1000000000000000000.00 |
大きな数も含むので、JavaScriptのNumberに変換して丸めるだけでは精度を保てません。負数や桁上がりも扱う必要があります。
問題文と出力上限を変えずに、2回ずつ依頼しました。
| モデル | 1回目 | 2回目 |
| DeepSeek V4.1 Flash | 本文が空・27.35秒 | 本文が空・24.06秒 |
| GPT-6 Astra | 16値すべて正解・9.16秒 | 16値すべて正解・8.72秒 |
GPTの回答は、返ってきたコードを実行して確かめました。DeepSeekは2回とも、APIが報告するcompletion tokensが設定上限の4,096に達し、本文は空でした。
これは「DeepSeekが丸め方を間違えた」という結果ではありません。今回のAPI経路と上限では、使える回答を受け取れなかったという結果です。空になった内部原因までは確認できていません。
もう1つは、空いている時間帯を求める関数
次に試したのは、区間の引き算です。予約可能な時間帯から、予約済みの時間帯を除く処理に相当します。
例として、次の区間を渡します。数字は時刻ではなく、単位を持たない整数です。
{
a: [[0, 10], [5, 20]], // 残したい範囲
b: [[3, 7], [9, 12], [18, 22]] // 取り除く範囲
}
期待する結果はこちらです。
[[0, 3], [7, 9], [12, 18]]
各区間は開始点を含み、終了点を含まない半開区間です。入力の順序はばらばらで、重なり、隣接、長さゼロの区間もあります。出力では空区間を除き、隣り合う区間をまとめます。
今回は基本5ケースに、固定乱数で作った100ケースを加えました。回答を見る前に正解を固定し、整数点の集合差と、別実装の境界点走査で一致を確認しています。
| モデル | 1回目 | 2回目 |
| DeepSeek V4.1 Flash | 105ケースすべて正解・17.53秒 | 本文が空・25.87秒 |
| GPT-6 Astra | 105ケースすべて正解・13.30秒 | 105ケースすべて正解・12.28秒 |
DeepSeekは最初の回答では、ちゃんと105ケースを通りました。2回目はcompletion tokensが4,096に達し、本文が空でした。GPTは2回とも通りました。
こちらは、小数の丸めと違って2回とも同じ差が出たわけではありません。少なくとも、同じ条件でも答えが返るかどうかに揺れがありました。2回だけなので、その発生率までは分かりません。
なお、105ケースは「105回AIに聞いた」という意味ではありません。1回の回答で作られた1つの関数に、105通りの入力を渡した結果です。
差が出なかったもの、GPT側で止まったものもある
差が出た例だけだと、比較の印象が偏ります。他の結果も残しています。
- 仕事の処理順を最適化する問題:両モデルとも正解。
- サイコロの条件付き確率:両モデルとも正解。
- JSON Pointerの実装:DeepSeekは初回が空回答、再確認では正解。GPTは2回とも正解。
- 費用と工数の2制約で案件を選ぶ問題、条件付き順列の問題:GPT側で、それぞれ約111秒後にHTTP 502。そこで実行を停止し、この2問はDeepSeekへの送信まで進みませんでした。
502は計算の誤答として採点していません。GPTというモデル自体、上流サービス、途中のゲートウェイのどこに原因があるかも、ここでは確定できていません。
最初に行った5モデル・12課題の60試行では、DeepSeekとGPTは全12課題で採点結果が一致していました。そのうち2課題は問題文に必要な項目名が抜けており、能力評価に使えないと後から分かりました。問題を作る側の不備も、モデルの失敗とは分けて扱う必要があります。
どう測ったか
今回の追加探索は、差が出る例を探すためのものです。広い用途を代表するランキングではありません。
- teaiのAPI経由で、指定した2モデルに送信。
- 出力上限は両方とも
max_tokens=4096、temperature=0、ストリーミングなし。temperatureが0でも、同一回答は保証されません。
- 外部ツールは使わせず、問題文だけを送信。テスト入力と正解は送っていません。
- 返されたコードは、固定したNodeイメージのネットワークなしDocker環境で採点。
- 秒数はリクエスト開始から応答を受け取るまでの時間。最初の文字が出る速さではありません。
- 同じトークン上限でも、モデル間で同じ計算量や同じ長さの回答を保証する条件ではありません。
今回の区間計算は4コールで、ゲートウェイの記録上は54クレジット、9円でした。先行試験と確認用コールを含む同じ専用キーの合計は363クレジット、60.50円です。上流の確定請求との照合は済んでいないため、これを確定した全費用とは呼んでいません。
予算上限は2,700円。失敗時の上流再試行も含む保守的な引当は通算2,674円となったため、追加比較はここで止めました。引当は実支出ではありません。
モデル選びで見たいことが1つ増えた
今回、内容の正解と誤答が分かれる例は見つかっていません。はっきり見えたのは、決めた上限内で使える答えが戻ってくるかどうかでした。
空き時間の問題では、DeepSeekが一度は全ケースを通しています。だから「このモデルには解けない」とは言えません。一方、仕事でそのまま使うなら、次の回に本文が空になることは無視できません。
Senteのモデル選びでも、正解率だけでなく、回答が完成したか、どれくらい待ったか、失敗した分まで含めていくらかかったかを見ていきます。今回の2問は、その違いが具体的に見えた例でした。
I asked AI for practical code. A correct answer once did not guarantee an answer next time.
Which model should we choose when using AI for work?
We are comparing short, practical tasks to help choose models in Sente. This time, rather than deciding in advance whether DeepSeek or GPT was better, we looked for problems where their results actually differed.
The API model names were DeepSeek V4.1 Flash (deepseek/deepseek-v4.1-flash) and GPT-6 Astra (openai/gpt-6-astra).
The difference we found was whether an answer came back at all. On one problem, the same model returned working code in one run and an empty answer in another.
First: decimal rounding
We asked for a JavaScript function that rounds a numeric string to two decimal places.
The rule was “at an exact tie, make the last retained digit even.” For example:
| Input | Expected |
1.005 | 1.00 |
1.015 | 1.02 |
-1.015 | -1.02 |
-0.005 | 0.00 |
999999999999999999.995 | 1000000000000000000.00 |
Because the inputs include large numbers, simply converting them to JavaScript Number and rounding cannot preserve precision. Negative numbers and carrying digits also need handling.
We made two requests to each model without changing the prompt or output limit.
| Model | Run 1 | Run 2 |
| DeepSeek V4.1 Flash | Empty answer · 27.35 s | Empty answer · 24.06 s |
| GPT-6 Astra | All 16 values correct · 9.16 s | All 16 values correct · 8.72 s |
We checked GPT’s answers by executing the returned code. In both DeepSeek runs, the API reported completion tokens reaching the configured limit of 4,096, with an empty answer body.
This does not mean “DeepSeek rounded incorrectly.” It means we did not receive a usable answer through this API route under this limit. The internal cause of the empty answers has not been verified.
Next: a function that finds available time ranges
Next we tried interval subtraction, analogous to removing booked time ranges from available ones.
Here is an example input. These numbers are unitless integers, not clock times.
{
a: [[0, 10], [5, 20]], // ranges to keep
b: [[3, 7], [9, 12], [18, 22]] // ranges to remove
}
The expected result is:
[[0, 3], [7, 9], [12, 18]]
Each interval is half-open: it includes its start and excludes its end. Inputs are unsorted and can overlap, touch or have zero length. Outputs must omit empty intervals and merge adjacent ones.
We added 100 seeded random cases to five basic cases. Expected answers were fixed before seeing model responses, and checked for agreement between integer-point set subtraction and an independently implemented boundary sweep.
| Model | Run 1 | Run 2 |
| DeepSeek V4.1 Flash | All 105 cases correct · 17.53 s | Empty answer · 25.87 s |
| GPT-6 Astra | All 105 cases correct · 13.30 s | All 105 cases correct · 12.28 s |
DeepSeek’s first answer passed all 105 cases. Its second reached 4,096 completion tokens with an empty answer body. GPT passed both times.
Unlike decimal rounding, the same difference did not appear in both runs. At least under these conditions, answer delivery varied. Two runs are not enough to determine the frequency.
Also, 105 cases does not mean we asked AI 105 times. It means we tested one function from one response against 105 different inputs.
Some tasks tied. Others stopped on the GPT side.
Showing only differences would give a skewed impression. We kept the other results too.
- Optimizing job processing order: both models were correct.
- Conditional probability with dice: both models were correct.
- Implementing JSON Pointer: DeepSeek’s first answer was empty; the confirmation run was correct. GPT was correct both times.
- Selecting projects under cost and effort constraints, and a constrained permutation problem: GPT requests each returned HTTP 502 after about 111 seconds. We stopped execution at those points and did not proceed to send those two tasks to DeepSeek.
We did not score 502 responses as wrong calculations. We have not established whether the cause lay in GPT itself, an upstream service or an intermediate gateway.
In the initial 60 attempts across five models and 12 tasks, DeepSeek and GPT had identical scores on all 12 tasks. We later found that two prompts omitted required field names, making those tasks unsuitable for evaluating ability. Defects in the tasks also need to be separated from model failures.
How we measured
This follow-up search was designed to find examples of differences. It is not a ranking representative of a broad range of uses.
- Requests were sent to the two specified models through teai’s API.
- Both used
max_tokens=4096, temperature=0, and no streaming. Zero temperature does not guarantee identical answers.
- Only the prompt was sent, with no external tools. Test inputs and expected answers were withheld.
- Returned code was graded in a network-disabled Docker environment with a pinned Node image.
- Times run from the start of the request to receipt of the response, not to the first character.
- An identical token cap does not guarantee equal computation or equal answer length across models.
The four interval-subtraction calls recorded 54 gateway credits, or JPY 9. The same dedicated key, including earlier tests and verification calls, recorded a total of 363 credits, or JPY 60.50. These have not been reconciled with settled upstream invoices, so we do not call them the confirmed total cost.
The budget ceiling was JPY 2,700. Conservative reservations, including upstream retries after failures, reached JPY 2,674, so we stopped further comparisons. Reservations are not actual spending.
One more thing to consider when choosing a model
We did not find a case where one model’s answer was correct and the other’s was wrong. What became clear was a difference in whether a usable answer came back within the chosen limit.
DeepSeek passed every case once on the availability problem, so we cannot say “this model cannot solve it.” But if we use it directly for work, an empty answer on the next run matters.
For model selection in Sente, we will look beyond correctness to whether the answer completed, how long we waited, and what it cost including failures. These two tasks made that distinction concrete.