Flagship models
Consider for demanding reasoning or complex evidence synthesis, with independent checks. No specific provider/model is verified or ranked by this launch.
Balanced models
Consider for bounded implementation, document work and structured tests when the relevant tools and validation are available.
Lightweight models
Consider for narrow classification, formatting or small well-specified steps. Validate the actual deliverable. Local execution is a deployment property, not a quality rank.
