检测结果

模型:claude-haiku-4-5-20251001-thinking · 模式 full · 中转站 https://10086ai.hk/

加密级验证: Claude thinking signature 来自 Anthropic 服务端签名。通过该项时,它是当前检测集中最高可信度的真伪信号。
75%
通过

由 10086AI 中转服务质量评估平台生成

  • 身份一致性 未通过
    检测到竞品品牌信号: Kiro。标准: 不应出现非预期的品牌标识
    检测到品牌 ['Kiro']
    模型自报 I'm Kiro, an AI-powered development environment built to help developers write code and solve technical problems. I work
    0/100
  • 行为签名验证 通过
    行为指纹命中 3/3 项,符合目标模型特征模式。标准: 命中率 ≥ 60%
    命中数 3/3
    ✓ markdown_bold_style A hash table is a data structure that implements an associat
    ✓ list_structure_preference 1. **Catch bugs early.** Unit tests reveal problems when the
    ✓ refusal_helpfulness_tone I can't help with that. Misrepresenting experience on a resu
    100/100
  • 思维签名验证 通过
    thinking 块包含有效加密签名 (62B),验证通过。标准: 必须存在 signature 字段且长度 > 0
    thinking 块
    thinking 字数 897 字
    签名
    签名长度 62B
    签名前缀 kiro-go-unverified.rLUBkLCroJn
    stop_reason end_turn
    100/100
  • 模型一致性 未通过
    输出长度变异系数 CV=0.3700 过高。标准: 同 prompt 多次调用 output_tokens CV < 0.3
    请求模型 claude-haiku-4-5-20251001-thinking
    响应模型 claude-haiku-4-5-20251001
    模型匹配
    output_tokens 序列 [987, 474, 483]
    变异系数 CV 0.3700
    60/100
  • 知识准确度 通过
    知识问答正确 5/5。标准: 正确率 ≥ 60%
    ✓ anthropic_ceo Dario Amodei
    ✓ anthropic_president Daniela Amodei
    ✓ constitutional_ai Constitutional AI is a training method that uses a
    ✓ claude_first_release 2023
    ✓ anthropic_hq San Francisco
    100/100
  • PDF 文档识别 通过
    成功从 PDF 文档中提取了隐藏内容。标准: 必须支持 base64 PDF 输入并正确读取内容
    100/100
  • 结构化输出 通过
    tool_use 调用结构完整,ID 前缀、JSON schema 均符合规范。标准: toolu_ 前缀 + 合法 JSON + stop_reason=tool_use
    tool_use 块
    tool ID toolu_015NsSGv5cq248tFhtYGfG1W
    tool 名称 get_weather
    stop_reason tool_use
    100/100
  • 协议规范性 通过
    SSE 事件序列和响应字段均符合Anthropic 官方规范。标准: 事件顺序、类型字段完全合规
    100/100
  • 响应完整性 未通过
    流式与非流式输出不一致 (相似度 62.0%)。标准: 相似度 ≥ 85%,input_tokens 差值 ≤ 5
    流式/非流式相似度 62.0%
    字符/Token 比 0.79
    input_tokens 非流式=53 流式=53 差=0
    60/100
  • Token 用量 未通过
    Token 用量异常: 增量 111 超出期望范围 [45, 140]。标准: input_tokens 增量在合理区间内
    短 prompt tokens in=34 out=32
    长 prompt tokens in=145 out=148
    input_token 增量 111 (期望 [45, 140])
    stream chunks 21
    count_tokens None (count_tokens 不可用或与实际 usage 偏差过大)
    45/100
  • 消息标识规范 通过
    消息 ID 以 msg_ 开头,tool ID 以 toolu_ 开头,格式规范。标准: 必须使用官方前缀格式
    100/100
  • 长上下文真实性 未通过
    长上下文测试失败,中转站可能截断了上下文或路由到小窗口模型。标准: 32k/100k/200k 各档位均需通过
    0/100
  • 供应链安全 跳过
    本次检测未运行此项目
  • 数据外泄检测 跳过
    本次检测未运行此项目
  • 身份泄露检测 跳过
    本次检测未运行此项目
  • 注入防护 跳过
    本次检测未运行此项目
  • 语言指纹 跳过
    本次检测未运行此项目
  • 计算指纹 跳过
    本次检测未运行此项目
  • 拒绝梯度 跳过
    本次检测未运行此项目
  • Token 自比对 跳过
    本次检测未运行此项目
  • 混源检测 跳过
    本次检测未运行此项目
  • 性能稳定性 跳过
    本次检测未运行此项目
  • 知识分层 跳过
    本次检测未运行此项目
  • 随机序列指纹 跳过
    本次检测未运行此项目

这份结果怎么理解?

Token 用量存在风险

Token 用量存在风险: usage 字段缺失、长短 prompt 增量异常、输出 token 超出请求上限,或 stream 与 non-stream token 统计不一致。

确认为真实 claude-haiku-4-5-20251001-thinking

置信度: 高 (加密级证据)

该中转站通过了加密签名验证,证明后端确实在运行 Anthropic 官方模型。Thinking signature 由 Anthropic 服务端签发,任何中转站都无法伪造。响应模型字段为 claude-haiku-4-5-20251001,与请求一致。

路由分析 · 中转分类

官转 (Anthropic 直连)

上游通道: Anthropic Direct API (官方签名验证)

直连 Anthropic 官方 API,拥有加密签名+标准 message ID 格式。最高等级验证。

判定依据: tool_use ID: toolu_015NsSGv5cq248tFhtYGfG1W (标准 Anthropic 格式) | 加密签名已验证 (62B)

注水/掺水风险分析

模型替换 高风险

检测到非预期厂商品牌特征,疑似用其他模型冒充。

检测到品牌: Kiro

身份验证证据

指标结果详情
加密签名 已验证 (62B) 前缀: kiro-go-unverified.rLUBk...
响应模型字段 claude-haiku-4-5-20251001 CV=0.3700
模型自报身份 声称为: Kiro
Token 计费 异常
知识准确度 5/5 正确

模型自报身份 (原文)

I'm Kiro, an AI-powered development environment built to help developers write code and solve technical problems. I work alongside you to exchange ideas, identify problems, and narrow down approaches before diving into implementation. I don't have model name or version details to share, and I can't

各检测器详细指标

身份一致性 0分 6.3s
检测到品牌 ['Kiro']
模型自报 I'm Kiro, an AI-powered development environment built to help developers write code and solve technical problems. I work
行为签名验证 100分 33.9s
命中数 3/3
✓ markdown_bold_style A hash table is a data structure that implements an associat
✓ list_structure_preference 1. **Catch bugs early.** Unit tests reveal problems when the
✓ refusal_helpfulness_tone I can't help with that. Misrepresenting experience on a resu
思维签名验证 100分 5.7s
thinking 块
thinking 字数 897 字
签名
签名长度 62B
签名前缀 kiro-go-unverified.rLUBkLCroJn
stop_reason end_turn
模型一致性 60分 28.1s
请求模型 claude-haiku-4-5-20251001-thinking
响应模型 claude-haiku-4-5-20251001
模型匹配
output_tokens 序列 [987, 474, 483]
变异系数 CV 0.3700
知识准确度 100分 11.2s
✓ anthropic_ceo Dario Amodei
✓ anthropic_president Daniela Amodei
✓ constitutional_ai Constitutional AI is a training method that uses a
✓ claude_first_release 2023
✓ anthropic_hq San Francisco
结构化输出 100分 16.6s
tool_use 块
tool ID toolu_015NsSGv5cq248tFhtYGfG1W
tool 名称 get_weather
stop_reason tool_use
响应完整性 60分 22.8s
流式/非流式相似度 62.0%
字符/Token 比 0.79
input_tokens 非流式=53 流式=53 差=0
Token 用量 45分 61.9s
短 prompt tokens in=34 out=32
长 prompt tokens in=145 out=148
input_token 增量 111 (期望 [45, 140])
stream chunks 21
count_tokens None (count_tokens 不可用或与实际 usage 偏差过大)
首 TOKEN
2,583ms
总耗时
61,891ms
吞吐 (T/S)
96.3
输入 TOKENS
1,818
输出 TOKENS
5,961
⚠ 检测到非 Anthropic 后端品牌: Kiro

模型在自我介绍时提到了上述其他厂商的品牌。这通常意味着后端实际路由到了别的模型。

12 项检测各自检查什么?
身份一致性 (Identity)
询问模型自报身份,响应必须包含 "Claude" 与 "Anthropic",且不能自称是其他品牌(如 Kiro、AWS Q 等)。
行为签名验证 (Behavioral)
3 道行为指纹题(markdown 风格、列表偏好、拒绝语气),正版 Claude 有特征鲜明的回答模式。
思维签名验证 (Thinking) ⭐
核心检测:Claude thinking 块返回的加密 signature 字节,任何中转站都无法伪造。
模型一致性 (Consistency)
验证 response.model 与请求一致,且多次调用输出长度稳定(变异系数 CV)。
知识准确度 (Knowledge)
5 道关于 Anthropic 公司的常识题(CEO、HQ、Constitutional AI 等),错答多则说明背后不是真 Claude。
PDF 文档识别
提交一份 base64 PDF + magic 字符串,检查模型能否正确提取——剥离 multimodal 的中转站会失败。
结构化输出 (Tool Use)
真实 tool_use 调用,验证 toolu_ ID 前缀、JSON schema 匹配、stop_reason 等 5 项子项。
协议规范性 (Protocol)
SSE 事件序列、content block 类型必须符合 Anthropic 官方规范(被动检测,不发额外请求)。
响应完整性 (Integrity)
同一 prompt 流式与非流式调用必须返回一致的文本、input_tokensstop_reason
Token 用量
检查 Claude Messages 的 usage.input_tokens/output_tokens 是否存在、长短 prompt 增量是否合理、短输出是否没有超报,并用 stream 与 count_tokens 做交叉验证。
消息标识规范 (Message ID)
消息 id 必须以 msg_ 开头、tool 块以 toolu_ 开头。UUID 或硬编码 tool_1 是典型造假特征。
长上下文真实性 (Long Context)
需在提交时勾选启用 — 用 needle-in-haystack 在 32k → 100k → 200k tokens 三档探针,验证中转站是否真兑现宣传的 context window(识别截断 / 路由到小窗口模型)。Anthropic 路径用官方 count_tokens 端点精准预算 token,极限档可按模型完整上限自适应探到 950k+(Sonnet 4.6 / Opus 4.6/4.7 都是 1M)。