発表された内容
Mistral AIは10月6日、Mistral Studio APIでMistral Large 4のpublic previewを開始し、10月末までにmodel weightを公開すると発表しました。同社は、coding、agent workflow、documentとvisual understanding、financeを含むprofessional work向けのnatively multimodalなmixture-of-experts modelと説明しています。発表によると、Mistralの欧州data centerにある3,800基のNVIDIA Grace Blackwell GPUで学習し、previewも同じinfrastructureから提供します。
model documentationはpreviewをversion v26.10とし、100万tokenのcontext window、structured output、function calling、document Q&A、batching、agent APIを掲載しています。発表は総parameter数を約1兆、active parameterを490億と説明する一方、現在のmodel pageは総1.05兆、active 520億、さらに16億parameterのvision encoderと記載しています。ここまではMistralが公表した仕様とbenchmark claimであり、以下はIneezaの本番運用分析です。
previewという名称はdeployment identityではない
Ineezaの分析: public previewは評価に有用ですが、本番workflowはproduct nameを不変の依存関係として扱えません。二つの公式pageにある丸め方の異なる仕様は、それ自体が信頼性問題ではありません。運用上のcontractには、正確なmodel identifier、provider region、API behavior、tokenizerとtool schema、evaluation set、観測日を結び付ける必要があるという注意点です。変化する「latest」aliasでは、その証跡を作れません。
既存modelを置き換える前に、代表的なtraceを固定したcandidateでreplayし、task成功率とrefusal behaviorを比較し、workflowごとのlatency、token使用量、tool-call shape、costを測定すべきです。promotionにはshadow trafficまたは上限付きcanaryを使い、明示的なrollback thresholdを設定します。publicなmodel nameが変わらなくても、preview更新は同じgateを再通過させる必要があります。
tool callingでも実行権限はapplication側に残る
Mistralのfunction-calling文書はmodel出力と実行を分離しています。modelはfunctionを選びargumentを生成しますが、functionの実行はdeveloperの責任です。また、tool callの強制、無効化、連続実行、並列実行に対応しています。この分離点が、資金移動、infrastructure変更、data公開を行うagentのcontrol pointです。
Ineezaの分析: 生成されたargumentは信頼できない提案として扱うべきです。実行層はside effectの前にbusiness policyで検証し、人またはserviceのintentを認証し、送信先と金額を制限し、budgetを適用し、idempotency keyを発行する必要があります。balance、nonce、inventory、approval stateを共有する操作ではparallel callは危険です。orchestratorは状態遷移を直列化するか、相互に影響しないことを証明すべきです。tool resultとretryはすべてdurableなexecution receiptへ結び付けます。
contextとmodalityの拡大は証跡境界も広げる
Ineezaの分析: 100万tokenのcontext windowとnative image understandingにより、filing、diagram、log、screenshotを一つのworkflowへ投入できます。しかしcontext容量はprovenanceではありません。金融または運用上の結論には、安定したsource identifier、content hash、pageまたはregion参照、抽出version、modelが実際に受け取った内容の記録が必要です。これがなければ、後のreviewerは一次証跡、OCR error、古いdocument、modelの要約を区別できません。
evaluation setには、正常なbenchmark入力だけでなく、敵対的なdocument、競合するversion、低品質scan、hidden instruction、欠落pageを含めるべきです。支払先、contract address、account value、approval termなど影響の大きい事実は、操作を認可する前にdeterministic parserまたは独立した検証経路を通す必要があります。
open weightは二つ目の運用対象を作る
Mistralは10月末までのweight公開を予告し、self-deploymentをmodel sovereigntyの価値として説明しています。ただし発表時点で利用できるartifactはhosted public-preview APIです。hosted previewと将来のself-hosted weightを同一のdeploymentとみなすべきではありません。
Ineezaの分析: migration計画ではhosted APIと将来のweight artifactを別々にqualifyし、modelとtokenizerのdigest、inference runtime、quantization、hardware topology、safety layer、generation settingを記録すべきです。outputとtool callのparity test、想定context長でのcapacity test、同じproviderやclusterに依存しないrollback経路が必要です。sovereigntyが配置と可用性の統制を改善するのは、組織がserving stack全体を再現、patch、監視、復旧できる場合に限られます。
Ineezaの見解
Mistral Large 4は、frontier規模のmultimodality、long context、agent tooling、欧州で提供するpreview、open-weight化の予告を一つにまとめた点で重要です。本番導入の安全な単位はmodel nameではありません。再現可能な評価、application側が所有する認可、source単位の証跡、統制されたpromotion、hosted環境とself-managed環境にまたがる検証済みrollbackを備えた、version管理されたdeployment contractです。