o - il@sddlmZddlZddlmZddlZddlmZddlm Z ddl m Z ddl m Z ddlmZmZdgZGd ddejZdS) ) annotationsN)Sequence)PatchEmbeddingBlockTransformerBlock)Conv)ensure_tuple_repis_sqrt ViTAutoEnccsBeZdZdZ         d$d%fd d! Zd"d#ZZS)&r a Vision Transformer (ViT), based on: "Dosovitskiy et al., An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale " Modified to also give same dimension outputs as the input size of the image  convF in_channelsintimg_sizeSequence[int] | int patch_size out_channels deconv_chns hidden_sizemlp_dim num_layers num_heads proj_typestr dropout_ratefloat spatial_dimsqkv_biasbool save_attnreturnNonec stt|std|dt|| |_t|| |_| |_t|j|jD]\}}||dkr