0
Article ? AI-assigned paper type based on the abstract. Classification may not be perfect — flag errors using the feedback button. Tier 2 ? Original research — experimental, observational, or case-control study. Direct primary evidence. Sign in to save

HCTSpeckle and Lightweight YUV Transformerwith ViT Encoder-Sandwich Decoder NetworkFor Underwater Microplastic High ResolutionImage Segmentation

International Journal of Image Graphics and Signal Processing 2026
Badugu Vimala Victoria, Kamil Reza Khondakar, Ravi Kumar Suggala

Summary

Scientists have developed a smart imaging system that can spot and sort microplastic pollution in murky underwater photos with impressive accuracy, about 97%. This matters because microplastics are turning up everywhere in our oceans and food chain, and better detection tools like this could help researchers track pollution more reliably, which is a key step toward understanding and reducing the health risks these tiny plastic particles may pose to people.

Microplastics are tiny particles made of plastic that are a significant source of pollution in the sea, and aredangerous to both the environment and human health. Nevertheless, the existing detection techniques are not robust andgeneral in dynamic underwater settings with varying microplastic shapes and sizes. To overcome these issues, a highresolution underwater microplastic segmentation framework was implemented that integrates HCTSpeckle and VisionTransformer Encoder–Sandwich Decoder Network. Initially, underwater sensors record continuous visual images. Thecaptured images are pre-processed with HCTSpeckle, a CNN-Transformer denoising network, for removing specklenoise, and it retains important structural information through the use of hybrid convolution-transformer blocks anddouble residual interactions. The denoised images were contrasted with the Lightweight YUV Transformer-basedNetwork, in which multistage squeeze-and-excitation fusion is used to improve the visibility of object boundaries. Theenhanced images are subjected to a hybrid Vision Transformer- Sandwich Decoder to produce precise underwatermicroplastic segmentation, where the ViT encoder captures high-level features that are globally correlated withoutdistorting positioning information in space by patch embeddings and self-attention mechanisms. They are decodedusing a Sandwich Decoder Network, which learns both local and global dependencies. Also, ranking and region-basedpooling priorities fine edges and microplastic structures, whereas the pixel-wise segmentation head preciselycategorizes the identified microplastics into fiber, film, pellet, and fragment types. The proposed approach attains thepixel accuracy of 0.97 and specificity of 0.98 with a Dice Coefficient of 0.934, which contains the effectivesegmentation result of underwater microplastic images. These findings ensure the framework efficiently identifies andcategorizes various types of microplastics in diverse underwater sceneries.

Share this paper