Fig 1.
The generic architecture for automatic captioning illustrates how the encoder merges images and vectors, and how the decoder generates sequence predictions.
Fig 2.
Captions from Flickr30k dataset in English 3rd row and Urdu 2nd row (ours).
Table 1.
Comparison of Urdu Image Captioning Studies.
Fig 3.
Image Captioning Taxonomy: An Overview of Corpora and Techniques.
Table 2.
English language corpora.
Table 3.
Non-English language corpus.
Table 4.
Manual Inspection and Correction.
Table 5.
UC-23-RY Corps Characteristics.
Fig 4.
Distribution of caption lengths (index vs text length).
Fig 5.
Flowchart of Urdu image caption generation process.
Table 6.
Division of training and testing images based on a split ratio.
Fig 6.
Proposed approach for Urdu image captioning.
Table 7.
Hyper-parameter of the trained model.
Table 8.
Time elapsed during training of ResNet-50-LSTM and NASNetLarge-LSTM model.
Fig 7.
Captions in Urdu using ResNet-50 and NASNetLarge models with their transliterations.
Fig 8.
ResNet-50 and NASNetLarge results for BLEU – (1,2,3,4) Scores.
Table 9.
BLEU scores for different datasets in different languages.
Table 10.
Feature and textural extraction model with random weights and urdu vector for a multimodal approach.
Table 11.
Parameter of Resnet-50 and NASNetLarge.