AUTO / OBSERVATION PILOT

仕事に使えるか、確かめる。

まずは、受付文を小さなJSONへ。答えの正しさだけでなく、使えるまでの修正・待ち時間・費用を記録します。モデルの採用更新は、実務の証拠が揃うまで保留です。

2026-09-21 · teai Engineering

JSON抽出パイロットを始める ↓

現在は、観測フェーズ。

合成テストだけによる自動採用は停止しました。promotion_mode は observe、promotions_active は false。同意を得た実務データと検証済みの受入結果が揃うまで、この状態を維持します。自己申告のフィードバックだけでは採用を解禁しません。

確認中 / Loading

合成テストの比較記録

失敗と未実行はバックエンドの集計をそのまま表示します。不明は「—」。合成テストの正答率は実務成功率ではなく、APIの表示は有料ライブ実行の成功証明ではありません。

答えから、仕事の完了まで。

01

用途を明示

UUIDと用途をヘッダーで指定。JSON抽出はGLM、規則・計算はKimiの基準モデルへ。Jevの分類呼び出しとその費用を省きます。

02

人が照合

元文書とJSONを照合し、accepted / rejected / needs_revisionを明示。サンプルが期待値に一致しても、自動で受入判定は送りません。

03

修正も含めて測る

同じUUIDで再試行し、試行数・時間・費用推計を積み上げます。acceptedでタスクを確定。台帳に生の入力文・出力文は保存しません。

対象はstream:false、空でない4,000文字以内のuserメッセージ1つ。履歴・system指示・ツール・画像は対象外です。300文字の下限はありません。正確なカタログの実モデルIDも指定でき、基準モデルとの比較に使えます。文章・コード全般の性能評価ではありません。

QUICK START / JSON EXTRACT

受付文を、確認できるJSONへ。

teaiに登録 → APIキーと残高を確認。以下は架空の申込受付です。curlかPythonの片方を選んでください。実行は通常のAPI料金が発生します。実務パイロットでは、利用に同意を得た受付文と事前に決めた正解・受入条件を用意します。

Register with teai → get an API key and check your balance. This fictional registration illustrates the flow. Choose either curl or Python; execution incurs normal API charges. For the real-work pilot, first obtain consent to use the notes and define reference answers and acceptance criteria.

export TEAI_API_KEY="te_your_api_key"

1. 明示的にタスクを登録して呼ぶ

curl
export TASK_ID="$(python3 -c 'import uuid; print(uuid.uuid4())')"
curl --fail-with-body --max-time 120 \
  -D auto-headers.txt -o auto-response.json \
  https://teai.io/v1/chat/completions \
  -H "Authorization: Bearer $TEAI_API_KEY" \
  -H "Content-Type: application/json" \
  -H "x-teai-task-id: $TASK_ID" \
  -H "x-teai-task-type: json_extract" \
  --data-binary @- <<'JSON'
{
  "model": "teai/auto",
  "messages": [{"role": "user", "content": "Extract this fictional event registration. Return only JSON with id, date, people; copy values without inference. Registration: R42; date: 2026-10-03; people: 3. No extra fields. Do not use Markdown code fences."}],
  "temperature": 0,
  "max_tokens": 256,
  "stream": false
}
JSON
python3 -m json.tool auto-response.json
python3 -c 'from pathlib import Path; print(Path("auto-headers.txt").read_text())'
export RECEIPT="$(python3 -c 'from pathlib import Path; h=Path("auto-headers.txt").read_text().splitlines(); print(next(v.split(":",1)[1].strip() for v in h if v.lower().startswith("x-teai-task-receipt:")))')"
curl --fail-with-body --max-time 30 \
  -H "Authorization: Bearer $TEAI_API_KEY" \
  "https://teai.io/api/v1/task-outcomes/$TASK_ID"

期待するJSON(実測結果ではありません)。finish_reasonがstopか、キー・型・値が元文書と一致するかを確認します。

{"id":"R42","date":"2026-10-03","people":3}

2. 結果と試行レシートを確認する

3. 照合後、人が判定する

業務で使えるならaccepted、使えなければrejected、修正が必要ならneeds_revisionを入力します。空欄は送信しません。最新のcompleteな試行だけを判定でき、accepted後のタスクは再試行できません。

curl:手動判定を送る
# Only after inspecting the source, output and GET summary above.
printf 'Decision (accepted/rejected/needs_revision; blank skips): '
read -r DECISION
export DECISION
case "$DECISION" in
  accepted|rejected|needs_revision)
    python3 -c 'import json,os; print(json.dumps({k:os.environ[v] for k,v in [("task_id","TASK_ID"),("receipt","RECEIPT"),("decision","DECISION")]}))' |
      curl --fail-with-body --max-time 30 \
        https://teai.io/api/v1/task-outcomes \
        -H "Authorization: Bearer $TEAI_API_KEY" \
        -H "Content-Type: application/json" --data-binary @-
    ;;
  *) printf 'No decision sent.\n' ;;
esac
Python:生成 → 確認 → 手動判定(curlの代わり)
python3 -m pip install openai

モデルがMarkdownコードフェンスを返す場合があります。以下は回答全体を囲む単一のフェンス(json指定または指定なし)だけを除去し、重複キー・NaNなどを拒否してJSONを解析します。文章中の波括弧からは抽出しません。プロンプトはモデル共通の出力形式を保証しません。

import json
import os
import re
import uuid
from urllib.request import Request, urlopen
from openai import OpenAI

def parse_json_answer(text):
    text = text.strip()
    fence = re.fullmatch(r'```(?:json)?\s*\n(.*?)\n```', text, re.S)
    if fence:
        text = fence.group(1)

    def reject_duplicates(pairs):
        result = {}
        for key, value in pairs:
            if key in result:
                raise ValueError(f"Duplicate key: {key}")
            result[key] = value
        return result

    def reject_constant(value):
        raise ValueError(f"Invalid JSON constant: {value}")

    return json.loads(text, object_pairs_hook=reject_duplicates,
                      parse_constant=reject_constant)

key = os.environ["TEAI_API_KEY"]
task_id = str(uuid.uuid4())  # Keep this ID for revisions of the same task.
client = OpenAI(api_key=key, base_url="https://teai.io/v1",
                max_retries=0, timeout=120.0)
prompt = (
    "Extract this fictional event registration. Return only JSON with "
    "id, date, people; copy values without inference. "
    "Registration: R42; date: 2026-10-03; people: 3. No extra fields. "
    "Do not use Markdown code fences."
)
print("Task ID:", task_id, flush=True)
raw = client.chat.completions.with_raw_response.create(
    model="teai/auto", messages=[{"role": "user", "content": prompt}],
    temperature=0, max_tokens=256, stream=False,
    extra_headers={"x-teai-task-id": task_id,
                   "x-teai-task-type": "json_extract"},
)
response = raw.parse()
receipt = raw.headers["x-teai-task-receipt"]
assert raw.headers["x-teai-task-id"] == task_id
print("Receipt:", receipt)
print("Model:", response.model, "Usage:", response.usage)
print("Reason:", raw.headers.get("x-teai-auto-reason", "unavailable"))
print("Source:", prompt)
choice = response.choices[0]
print("Output:", choice.message.content)
if choice.finish_reason != "stop":
    raise RuntimeError(f"Incomplete answer: {choice.finish_reason}")
data = parse_json_answer(choice.message.content or "")
if not isinstance(data, dict) or set(data) != {"id", "date", "people"}:
    raise ValueError("Unexpected fields; inspect before retrying")
if (type(data["people"]) is not int or data["people"] < 1
        or not all(isinstance(data[k], str) for k in ("id", "date"))):
    raise ValueError("Invalid field types")

def outcomes(path, body=None):
    payload = None if body is None else json.dumps(body).encode()
    req = Request("https://teai.io/api/v1/task-outcomes" + path,
                  data=payload, headers={"Authorization": f"Bearer {key}",
                                        "Content-Type": "application/json"})
    with urlopen(req, timeout=30) as result:
        return json.load(result)

summary = outcomes("/" + task_id)
print(json.dumps(summary, ensure_ascii=False, indent=2))
if summary["current_receipt"] != receipt:
    raise RuntimeError("A newer attempt exists; inspect it first")
# A schema check or sample match is NOT acceptance. Review the source yourself.
decision = input("After review: accepted/rejected/needs_revision (blank skips): ").strip()
if decision in {"accepted", "rejected", "needs_revision"}:
    result = outcomes("", {"task_id": task_id, "receipt": receipt,
                           "decision": decision})
    print(json.dumps(result, ensure_ascii=False, indent=2))
else:
    print("No decision sent.")

修正時はTASK_IDを作り直さず、同じ用途で再実行します。新しい試行の費用・時間も加算されます。新規の業務は新しいUUIDへ。実モデルの比較は同じ入力・受入条件と別の比較用UUIDで行い、初回回答だけでなく修正を含む合計を比べます。

台帳の読み方・エラー時の確認
  • attempt_countは総試行数、correction_countは追加試行数。accepted_first_timeは初回受入ならtrue、再試行後ならfalse、未受入ならnullです。
  • total_attempt_elapsed_msは各試行時間の合計。elapsed_overall_msは修正間隔も含む最初の開始から最後の終了まで。未観測・未完了はnullのままです。
  • elapsed_until_decision_msは最初の開始から最新試行の自己申告判定まで(確認時間を含む)。未判定時や再試行後の判定前はnullです。
  • unknown_cost_attemptsは費用不明の試行数。feedback_source=self_reported、cost_basis=estimate_not_billed、verified_corpus=false、production_evidence_eligible=falseです。
  • 400はヘッダー・入力条件、401は認証、402は残高を確認。404は所有者・UUID・レシートを確認。409は古いレシート、未完了、判定済み、accepted後の再試行など。429/5xxや応答不明時はGETで台帳を確認してから判断します。

次の採用判断に必要なもの。

2026-09-21 本番確認:架空の受付文でcurl・Pythonから正しいJSONを取得し、レシート照会、要修正判定の冪等性、2回の試行の合計記録など15項目を確認しました。2回の応答時間合計は3,304ms、台帳の費用推計合計は0.003455406円でした。小規模な接続確認であり、実務の成功率・請求額・一般的な速度を示すものではありません。

同意済みの実務コーパス、事前に定めた正解と受入条件、基準モデルとの比較、独立した検証が必要です。合成テスト24件や自己申告のacceptedだけで、実務性能の証明や自動採用へ進めることはありません。

初回の確認では正しいJSONがコードフェンスで囲まれ、直接のJSON解析が失敗しました。回答全体が単一のJSONコード枠である場合だけ外す処理へ修正し、再検証しています。人による実務の受入判定は別に必要です。

運営の評価予算と履歴

運営の評価予算はUTC月3,000円・日300円。呼び出す前に費用枠を予約し、不明な呼び出しの予約は残します。予約も費用推計も実請求ではなく、お客様のAPI料金とは別です。以下は過去の採用・復帰イベントで、現在の有効設定を示すものではありません。

読み込み中 / Loading