Date of Award
5-2026
Document Type
Dissertation
Publisher
Santa Clara : Santa Clara University, 2026
Degree Name
Doctor of Philosophy (PhD)
Department
Computer Science and Engineering
First Advisor
Ying Liu
Abstract
Driven by rapid progress in machine learning together with innovations in modern imaging and transmission technologies, visual data now primarily undergoes machine-oriented processing, which has in turn spurred the emergence of image coding for machines (ICM). After the codec completes the compression and delivery of the visual data, the resulting representation is subsequently passed to the machine vision task networks for further processing. These downstream vision modules include a broad range of tasks, such as image classification, semantic segmentation, object detection and instance segmentation. Furthermore, after the ICM system has fulfilled the requirements of visual tasks, there is a desire to enable it to perform complete image reconstruction, to meet the needs of human visual perception.
We focus on using neural networks to build ICM frameworks, with the goal of achieving end-to-end trainability. To build the learned ICM frameworks capable of handling machine tasks and/or image reconstruction that require image fidelity, we have progressively proposed and implemented the following three proposals.
- In order to use less bitrates while dealing with image classification task, we proposed a deep learning-based image compression network that combines the side information that assists the main stream (latent information) encoding with the image classification vi network to form the side information driven image coding for machines (SIIC) framework. Because the hyperprior—serving as side information—encodes higherlevel semantic cues while requiring only a small share of the total bitrate, the proposed approach is able to substantially reduce transmission bandwidth yet still sustain strong image-classification performance. [28]
- In order to handle machine tasks (image classification, semantic segmentation) and image reconstruction for human eye, we proposed a scalable hybrid machine-human vision ICM framework SICMH. Its base layer supports semantic segmentation, image classification, and coarse-level image reconstruction, whereas the enhancement layer is dedicated to producing fine-grained reconstruction results. In the proposed SICMH framework, semantic cues encoded in the hyperpriors are exploited through a multiscale feature fusion module, enabling more effective extraction of high-level representations and consequently yielding improvements in semantic segmentation mIoU, image classification accuracy, and coarse reconstruction PSNR [72]. For finegrained image reconstruction, SICMH combines the side information conveyed by the base layer with the residual signals supplied by the enhancement layer, allowing the restored images to better conform to human visual perception.
In order to jointly address object detection, instance segmentation, and image compression tasks within a unified framework, we propose a two-hyperprior multitask image coding for machines system, named TMFBC-SP, which incorporates a Blur CLIP context model (BCcm) and a ResNet Pyramid Hierarchical Feature Extractor (RPHFE) [30]. The BCcm module leverages CLIP’s global semantic understanding and PixelCNN’s spatial dependency modeling to predict more accurate latent distributions, thereby reducing bitrate consumption. Meanwhile, the RPHFE module utilizes multi-scale features generated through downsample and upsample resize pairs to enhance detection and segmentation accuracy. TMFBC-SP achieves superior performance compared to state-of-the-art methods vii in terms of mean Average Precision (mAP) for both detection and segmentation, while using significantly fewer bitrates. Additionally, for image reconstruction, our method outperforms traditional codecs such as BPG and learned codecs such as coarse-to-fine frameworks by achieving higher PSNR under the same or lower bitrate settings.
Recommended Citation
Zhang, Zhongpeng, "Learned Hybird Image Coding: From Machine Tasks to Pixel Fidelity" (2026). Engineering Ph.D. Theses. 69.
https://scholarcommons.scu.edu/eng_phd_theses/69
