<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Transformer on 标准答案</title><link>https://www.xin800.com/tags/transformer/</link><description>Recent content in Transformer on 标准答案</description><generator>Hugo</generator><language>zh-CN</language><lastBuildDate>Wed, 05 Aug 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://www.xin800.com/tags/transformer/index.xml" rel="self" type="application/rss+xml"/><item><title>Transformer 原始论文为什么N=6层（Encoder6层、Decoder6层）</title><link>https://www.xin800.com/post/why-transformer-6-layers/</link><pubDate>Wed, 05 Aug 2026 00:00:00 +0000</pubDate><guid>https://www.xin800.com/post/why-transformer-6-layers/</guid><description>业务通俗比喻 把每一层Block想象成一道加工工序： - 第1‑2层：看懂字面、词语、局部语法（看懂相邻词语关系） - 第3‑4层：理解句子句法、短语关系 - 第5‑6层：提炼全局深层语义、长距离逻辑（跨很远的词的指代、逻辑） 句子（张量）一层层流过每道工序，逐层把原始输入打磨成高级语义特征。 👉</description></item><item><title>Transformer 注意力整套完整流程梳理</title><link>https://www.xin800.com/post/transformer-attention-pipeline/</link><pubDate>Fri, 31 Jul 2026 00:00:00 +0000</pubDate><guid>https://www.xin800.com/post/transformer-attention-pipeline/</guid><description>从向量到QKV到点积到缩放到Softmax到加权融合，完整的Transformer注意力流程，全程业务语言，供应链案例贯穿。</description></item><item><title>什么是交叉注意力机制（Cross-Attention）</title><link>https://www.xin800.com/post/what-is-cross-attention/</link><pubDate>Fri, 31 Jul 2026 00:00:00 +0000</pubDate><guid>https://www.xin800.com/post/what-is-cross-attention/</guid><description>交叉注意力机制实现两组不同业务信息之间的关联匹配，Q来自A数据集，K/V来自B数据集，是机器翻译和知识库匹配的核心技术。</description></item><item><title>什么是自注意力机制（Self-Attention）</title><link>https://www.xin800.com/post/what-is-self-attention/</link><pubDate>Fri, 31 Jul 2026 00:00:00 +0000</pubDate><guid>https://www.xin800.com/post/what-is-self-attention/</guid><description>自注意力机制实现同一组数据内部所有条目的两两关联计算，依靠QKV向量与相似度打分自动识别远距离业务关联，是Transformer的核心基础。</description></item><item><title>多头注意力机制（Multi-Headed Self-Attention）</title><link>https://www.xin800.com/post/multi-head-attention/</link><pubDate>Fri, 31 Jul 2026 00:00:00 +0000</pubDate><guid>https://www.xin800.com/post/multi-head-attention/</guid><description>多头注意力赋予注意力多种子表达方式，8组独立QKV矩阵并行分析，同时捕捉因果、时序、实体、动作等多种关系。</description></item><item><title>点积与余弦相似度</title><link>https://www.xin800.com/post/dot-product-vs-cosine/</link><pubDate>Fri, 31 Jul 2026 00:00:00 +0000</pubDate><guid>https://www.xin800.com/post/dot-product-vs-cosine/</guid><description>通过Q-K点积自动算出任意两条业务信息之间关联程度，详解Thing、query-vector、key-vec等概念，区分点积与余弦相似度的本质差异。</description></item><item><title>点积业务场景与计算逻辑</title><link>https://www.xin800.com/post/dot-product-business/</link><pubDate>Fri, 31 Jul 2026 00:00:00 +0000</pubDate><guid>https://www.xin800.com/post/dot-product-business/</guid><description>点积是一套快速计算公式，用来衡量两组向量大方向是否一致，输出一个数字代表匹配程度，是注意力机制计算关联程度的基础运算。</description></item></channel></rss>