Multi-Task Deep Learning for Damage Classification, Structured Caption Generation, and Image Synthesis of RC Beam-Column Joint Failures
| dc.contributor.advisor | Lee, Min-Ho | |
| dc.contributor.author | Saparbekov, Yerassyl | |
| dc.date.accessioned | 2026-05-26T11:41:52Z | |
| dc.date.issued | 2026-04-11 | |
| dc.description.abstract | Reinforced concrete beam-column joints are among the most structurally critical and damage-prone elements in framed buildings, yet the visual documentation of their failure patterns still relies overwhelmingly on manual expert inspection. This thesis investigates whether deep learning can automate three complementary aspects of that documentation process, damage classification, structured caption generation, and synthetic image augmentation, when only a small, laboratory-scale image dataset is available. We assemble a purpose-built corpus of 234 images depicting beam (B), beam-joint (BJ), and joint (J) failure modes, each annotated with a structured seven-field caption that encodes element type, material, damage type, location, severity, failure mechanism, and a free-text observer description. The dataset is small (163 training, 71 test images) and domain-specific, which places it outside the operating assumptions of most contemporary vision-language models. The work is organised around three tasks. Task 1 benchmarks five convolutional and transformer-based classifiers under both pretrained and scratch conditions, paired with Grad-CAM++ interpretability analysis to verify whether high accuracy corresponds to structurally meaningful attention. Task 2 evaluates sixteen encoder-decoder captioning architectures across thirty-two training runs, establishing performance ceilings for the pre-2023 captioning paradigm on structured output. Task 3 extends the benchmark to recent large vision-language models (2024–2025), comparing fine-tuned and zero-resource configurations to determine the conditions under which modern LLMs can produce domain-specific structured captions. Among the principal findings, Swin Transformer pretrained on ImageNet achieves the highest classification accuracy (81.69%) and is the only model whose structural attention is statistically significant (Class Activation Correlation p = 0.013). For captioning, fine-tuned Gemma-3-4B attains the best BLEU-4 score (0.4494), but models below approximately one billion parameters fail entirely regardless of fine-tuning. Architecture generation proves more predictive of performance than raw parameter count, and the encoder-decoder paradigm is shown to have reached a practical ceiling by 2022. Building on these benchmarks, a physics-informed multi-task SegFormer-B2 pipeline that augments RGB input with six structural map channels and an auxiliary clDice segmentation head reaches 76.06% test accuracy, with the structural maps contributing +11.27 percentage points over an RGB-only baseline. A complementary label-conditioned captioning experiment, in which classifier-predicted damage class labels are injected as text prefixes into a LoRA fine-tuned Gemma-3-4B, recovers 82.1% of the METEOR gain that would be achievable with oracle labels, demonstrating that an imperfect classifier can still transfer meaningful structural priors to the captioner. These results offer guidance for deploying vision-language pipelines under small-data, structured-output constraints in civil engineering inspection contexts. | |
| dc.identifier.citation | Saparbekov, Y. (2026). Multi-Task Deep Learning for Damage Classification, Structured Caption Generation, and Image Synthesis of RC Beam-Column Joint Failures. Nazarbayev University School of Engineering and Digital Sciences | |
| dc.identifier.uri | https://nur.nu.edu.kz/handle/123456789/18741 | |
| dc.language.iso | en | |
| dc.publisher | Nazarbayev University School of Engineering and Digital Sciences | |
| dc.rights | Attribution-NonCommercial 3.0 United States | en |
| dc.rights.uri | http://creativecommons.org/licenses/by-nc/3.0/us/ | |
| dc.subject | Computer Vision | |
| dc.subject | Deep Learning | |
| dc.subject | Machine Learning | |
| dc.subject | Computer Science | |
| dc.subject | Image Captioning | |
| dc.subject | Damage Classification | |
| dc.subject | Reinforced concrete (RC) structures | |
| dc.subject | Beam-column joints | |
| dc.subject | Structural damage assessment | |
| dc.subject | Structured caption generation | |
| dc.title | Multi-Task Deep Learning for Damage Classification, Structured Caption Generation, and Image Synthesis of RC Beam-Column Joint Failures | |
| dc.type | Master`s thesis |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- Thesis_Yerassyl_Saparbekov.pdf
- Size:
- 13.26 MB
- Format:
- Adobe Portable Document Format
- Description:
- Master`s thesis