<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Attention-Maps | Laurent Perrinet</title><link>https://laurentperrinet.github.io/tag/attention-maps/</link><atom:link href="https://laurentperrinet.github.io/tag/attention-maps/index.xml" rel="self" type="application/rss+xml"/><description>Attention-Maps</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en</language><copyright>This material is presented to ensure timely dissemination of scholarly and technical work. Copyright and all rights therein are retained by authors or by other copyright holders. All persons copying this information are expected to adhere to the terms and constraints invoked by each author's copyright. In most cases, these works may not be reposted without the explicit permission of the copyright holder. This work is licensed under a Creative Commons Attribution-Noncommercial-Share Alike 3.0 Unported License Please note that multiple distribution, publication or commercial usage of copyrighted papers included in this website would require submission of a permission request addressed to the journal in which the paper appeared.</copyright><lastBuildDate>Thu, 01 Jan 2026 00:00:00 +0000</lastBuildDate><image><url>https://laurentperrinet.github.io/media/icon_hu_f2990a9a83ba401.png</url><title>Attention-Maps</title><link>https://laurentperrinet.github.io/tag/attention-maps/</link></image><item><title>A saccade-inspired approach to image classification using vision transformer attention maps</title><link>https://laurentperrinet.github.io/publication/dallain-26/</link><pubDate>Thu, 01 Jan 2026 00:00:00 +0000</pubDate><guid>https://laurentperrinet.github.io/publication/dallain-26/</guid><description>
&lt;figure id="figure-saccade-selection-method-a-the-input-image-of-dimensionh-wis-split-intoh16wnsized-patches-and-embeddedinto-token-vectors-b-the-tokens-are-passed-through-the-dino-transformer-and-attention-flow-from-patch-tokens-to-clstoken-white-arrows-are-extracted-and-reshaped-into-one-attention-map-per-attention-head-c-the-multiple-attention-maps-arefused-into-one-by-taking-the-maximum-value-across-heads-d-the-highest-attention-locations-define-square-regionssaccades-whose-tokens-are-retained-e-selected-regions-are-revealed-sequentially-and-the-image-variants-are-classified-by-a-pre-trained-linear-head"&gt;
&lt;div class="d-flex justify-content-center"&gt;
&lt;div class="w-100" &gt;&lt;img alt="Saccade selection method: (a.) The input image of dimensionH× Wis split intoH16×Wnsized patches and embeddedinto token vectors. (b.) The tokens are passed through the DINO transformer, and attention flow from patch tokens to [CLS]token (white arrows) are extracted and reshaped into one attention map per attention-head. (c.) The multiple attention maps arefused into one by taking the maximum value across heads. (d.) The highest-attention locations define square regions(“saccades”) whose tokens are retained. (e.) Selected regions are revealed sequentially, and the image variants are classified by a pre-trained linear head." srcset="
/publication/dallain-26/saccade_selection_hu_a552a3a1b7e1914d.webp 400w,
/publication/dallain-26/saccade_selection_hu_d696ecf3b868fbc1.webp 760w,
/publication/dallain-26/saccade_selection_hu_8e4f67b6cc46341f.webp 1200w"
src="https://laurentperrinet.github.io/publication/dallain-26/saccade_selection_hu_a552a3a1b7e1914d.webp"
width="760"
height="399"
loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;figcaption&gt;
Saccade selection method: (a.) The input image of dimensionH× Wis split intoH16×Wnsized patches and embeddedinto token vectors. (b.) The tokens are passed through the DINO transformer, and attention flow from patch tokens to [CLS]token (white arrows) are extracted and reshaped into one attention map per attention-head. (c.) The multiple attention maps arefused into one by taking the maximum value across heads. (d.) The highest-attention locations define square regions(“saccades”) whose tokens are retained. (e.) Selected regions are revealed sequentially, and the image variants are classified by a pre-trained linear head.
&lt;/figcaption&gt;&lt;/figure&gt;</description></item></channel></rss>