Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

< Back to Article

Fig 1.

The generic architecture for automatic captioning illustrates how the encoder merges images and vectors, and how the decoder generates sequence predictions.

More »

Fig 1 Expand

Fig 2.

Captions from Flickr30k dataset in English 3rd row and Urdu 2nd row (ours).

More »

Fig 2 Expand

Table 1.

Comparison of Urdu Image Captioning Studies.

More »

Table 1 Expand

Fig 3.

Image Captioning Taxonomy: An Overview of Corpora and Techniques.

More »

Fig 3 Expand

Table 2.

English language corpora.

More »

Table 2 Expand

Table 3.

Non-English language corpus.

More »

Table 3 Expand

Table 4.

Manual Inspection and Correction.

More »

Table 4 Expand

Table 5.

UC-23-RY Corps Characteristics.

More »

Table 5 Expand

Fig 4.

Distribution of caption lengths (index vs text length).

More »

Fig 4 Expand

Fig 5.

Flowchart of Urdu image caption generation process.

More »

Fig 5 Expand

Table 6.

Division of training and testing images based on a split ratio.

More »

Table 6 Expand

Fig 6.

Proposed approach for Urdu image captioning.

More »

Fig 6 Expand

Table 7.

Hyper-parameter of the trained model.

More »

Table 7 Expand

Table 8.

Time elapsed during training of ResNet-50-LSTM and NASNetLarge-LSTM model.

More »

Table 8 Expand

Fig 7.

Captions in Urdu using ResNet-50 and NASNetLarge models with their transliterations.

More »

Fig 7 Expand

Fig 8.

ResNet-50 and NASNetLarge results for BLEU – (1,2,3,4) Scores.

More »

Fig 8 Expand

Table 9.

BLEU scores for different datasets in different languages.

More »

Table 9 Expand

Table 10.

Feature and textural extraction model with random weights and urdu vector for a multimodal approach.

More »

Table 10 Expand

Table 11.

Parameter of Resnet-50 and NASNetLarge.

More »

Table 11 Expand