Limited-Time Offer: Enjoy 50% Savings! Ends in 00h 00m 00s Coupon code: 50OFF
Skip to content

Free NVIDIA Generative AI Multimodal NCA-GENM Exam Questions

Page: 1 / 6 Total 56 questions

Want more questions? Get Premium Access.

Question 1

You are developing a GenAI-Multimodal system that uses data from various sources. What is one potential issue you need to consider in relation to bias in data?

Correct Answer: A. The data used to train the AI system may not be representative of the population it is intended to serve.
Explanation:

Representativeness bias occurs when a training dataset systematically over- or under-samples subpopulations relative to the population the deployed system will actually encounter --- for example, a facial recognition dataset skewed toward lighter-skinned faces, or a multimodal medical dataset drawn predominantly from one demographic group. Because models learn statistical patterns from their training distribution, an unrepresentative dataset produces a model whose accuracy, calibration, and fairness properties degrade for underrepresented groups, even when aggregate accuracy metrics look acceptable.

This is precisely why aggregate accuracy is an insufficient safeguard: option B's framing --- that bias doesn't matter 'as long as predictions are accurate' --- conflates overall accuracy with subgroup accuracy, and a model can post strong aggregate numbers while systematically failing specific populations. Option D is factually false; AI systems have no inherent neutrality --- they inherit and can amplify whatever patterns (including societal biases) exist in their training data and objective function. Option C is also incorrect: mitigating representativeness bias is significantly cheaper and more effective when addressed at the data-collection and curation stage --- through stratified sampling, bias audits, and diverse data sourcing --- than after deployment, when it becomes a retraining and remediation problem, and by then real-world harm may have already occurred.


Question 2

What is a main application of Triton Inference Server?

Correct Answer: D. Triton Server can be used to deploy neural networks from various frameworks.
Explanation:

NVIDIA Triton Inference Server is an open-source model-serving platform designed to standardize production deployment across heterogeneous model formats and hardware backends. Its defining capability is multi-framework support: a single Triton instance can concurrently serve models trained in TensorFlow, PyTorch, ONNX Runtime, TensorRT, OpenVINO, and custom Python/C++ backends, exposing them through unified HTTP/REST and gRPC inference APIs. This eliminates the need for framework-specific serving stacks and lets teams standardize deployment infrastructure independent of how a given model was trained.

Option B is a common but incorrect assumption --- Triton explicitly supports both GPU and CPU inference, which matters for cost-sensitive or edge deployments where GPU availability is limited. Option C confuses Triton with cuGraph, a separate RAPIDS library for GPU-accelerated graph analytics (unrelated to model serving), and option A describes a generative task (denoising diffusion), which is a *model capability*, not a Triton *server* function --- Triton can host such a model, but 'generating images from noise' is not what the server itself does.

Triton also provides dynamic batching, concurrent model execution, model ensembling (chaining pre/post-processing with inference), and metrics export --- features tested elsewhere in the Software Development and Engineering and Performance Optimization domains.


Question 3

Which visualization technique is suitable for representing the distribution of performance scores for different multimodal ML models over different modalities?

Correct Answer: C. Box plot
Explanation:

A box plot (box-and-whisker plot) summarizes the distribution of a numeric variable --- median, interquartile range, and outliers --- as a single compact glyph, and critically, multiple box plots can be placed side by side to compare distributions across categorical groupings. This makes it well suited to the scenario described: comparing the spread and central tendency of performance scores across several models, further faceted by modality, in one readable figure. Box plots make skew, variance, and outlier prevalence immediately comparable across groups in a way a single summary statistic (like mean accuracy) cannot.

A histogram (B) shows the distribution of a single variable well but does not scale cleanly to side-by-side comparison across many model/modality combinations without becoming visually cluttered. A heatmap (A) is excellent for showing a matrix of values (e.g., mean score per model modality pair) but represents point estimates, not distributions --- it cannot convey variance or spread. A pie chart (D) is inappropriate for any continuous performance metric.

In practice, a violin plot --- which overlays a kernel density estimate on the box plot's summary statistics --- is often preferred when the underlying distribution's shape (e.g., bimodality) matters, but among the given options, the box plot is the correct choice for distributional comparison across groups.


Question 4

During the process of data cleansing, which of the following steps is NOT typically performed?

Correct Answer: C. Collecting additional data
Explanation:

Data cleansing (or data cleaning) operates on data you already have: it identifies and resolves quality issues within an existing dataset --- handling missing values (A), removing duplicate records (D), correcting formatting or type inconsistencies (B), fixing structural errors, and standardizing units or encodings. Collecting additional data (C) belongs to a conceptually earlier and separate phase of the pipeline: data acquisition or data collection, which determines what data enters the pipeline in the first place, rather than what is done to improve the quality of data already collected.

This distinction matters operationally: a cleansing step is typically deterministic and reversible against the existing dataset (you can inspect, log, and audit exactly which rows were dropped or imputed), whereas collecting more data is a scoping decision that may require new labeling budgets, new consent/privacy review, or new data-source integration --- a materially different workflow with different stakeholders.

That said, insufficient data volume discovered *during* cleansing (e.g., after removing corrupted records the sample size becomes too small for the target class) can trigger a decision to go back and collect more --- but that action itself is not classified as a cleansing step; it is the trigger for restarting an earlier pipeline stage.


Question 5

Which metric is commonly used to evaluate machine-translation models?

Correct Answer: D. BLEU score
Explanation:

BLEU (Bilingual Evaluation Understudy) is the standard automatic metric for evaluating machine translation quality. It measures n-gram precision --- the overlap of contiguous word sequences (unigrams through typically 4-grams) between the model's translated output and one or more human reference translations --- combined with a brevity penalty to discourage overly short translations that could otherwise achieve artificially high precision. BLEU scores range from 0 to 1 (or 0-100 as a percentage), with higher scores indicating closer alignment to reference translations.

The distractors represent metrics standard to other task families: F1 score (A) evaluates classification tasks by balancing precision and recall over discrete positive/negative predictions, ill-suited to open-ended text generation where there is no fixed set of 'correct' tokens. Accuracy (B) similarly assumes a discrete correct/incorrect judgment, inappropriate for translation where multiple valid phrasings can convey the same meaning. Mean Absolute Error (C) is a regression metric measuring average magnitude of numeric prediction error, irrelevant to text output evaluation entirely.

It's worth noting BLEU has known limitations --- it correlates imperfectly with human judgments of fluency and can penalize valid paraphrases --- which has motivated complementary metrics like METEOR, ROUGE (more common for summarization), and learned metrics like BERTScore, though BLEU remains the benchmark most commonly referenced for translation specifically.


Question 6

You are working with a large dataset and want to visualize the distribution of a continuous variable. Which type of data visualization would be most appropriate?

Correct Answer: A. Histogram chart
Explanation:

A histogram bins a continuous variable into contiguous intervals and plots the frequency (or density) of observations falling into each bin, making it the standard tool for visualizing the shape of a continuous distribution --- skewness, modality, spread, and outliers are all immediately visible. This distinguishes it from a bar chart (B), which is designed for discrete or categorical variables where bars are separated and ordering is often arbitrary; applying a bar chart to continuous data loses the notion of a numeric scale between categories.

A line chart (C) is appropriate for showing trends of a variable across an ordered sequence, typically time, not for summarizing the overall shape of a value distribution. A pie chart (D) shows proportions of a whole across categorical segments and becomes visually unreadable and statistically meaningless for continuous data with many possible values.

In practice, histogram bin width is a critical hyperparameter: too few bins oversmooth the distribution and hide multimodality, while too many bins introduce noise. Tools like Freedman-Diaconis or Sturges' rule provide principled starting points, and kernel density estimates (KDE) are often overlaid as a smoothed alternative when bin-width sensitivity is a concern.


Question 7

In the development of Trustworthy AI, what is the significance of 'Certification' as a principle?

Correct Answer: D. It involves verifying that AI models are fit for their intended purpose according to regional or industry-specific standards.
Explanation:

Within Trustworthy AI frameworks, 'Certification' is best understood as the formal verification process confirming that an AI system meets defined standards of fitness-for-purpose --- whether those standards are set by regulatory bodies, industry consortia, or internal governance frameworks --- for the specific context in which the system will be deployed. This is distinct from, though related to, the broader Trustworthy AI principles of ethics (option A), transparency (option B), and legal compliance (option C): certification is the *verification mechanism* that attests a system satisfies applicable standards, rather than being one of those underlying values itself.

The distinction between C and D is subtle and worth being precise about: C describes compliance as an obligation ('must follow laws and regulations'), while D describes certification as a verification activity ('confirming fitness according to standards') --- certification is the audit/attestation process, and compliance is one of the things that process may confirm. A system can be legally compliant without having undergone formal certification, and certification processes often assess criteria broader than legal compliance alone, including performance benchmarks, robustness testing, and domain-appropriate validation (e.g., clinical validation standards for a medical imaging model).

In practice, certification connects Trustworthy AI to concrete deployment gates: a healthcare AI model, for instance, may require certification against medical device standards before clinical use --- the verification step, not merely the legal requirement, is the 'Certification' principle's substance.


Question 8

What is the purpose of a kernel in a Convolutional Neural Network (CNN)?

Correct Answer: A. To perform convolution operations on input data.
Explanation:

A kernel (or filter) in a CNN is a small matrix of learnable weights that slides across the input (an image, feature map, or intermediate activation) computing a dot product at each spatial position --- the convolution operation. Each kernel is trained to detect a specific local pattern: early-layer kernels typically learn to detect low-level features like edges and color gradients, while kernels in deeper layers combine these into detectors for more complex, higher-level patterns (textures, object parts, and eventually whole-object representations as receptive fields grow with depth). A convolutional layer typically applies many kernels in parallel, each producing its own output channel, collectively forming the layer's feature map.

The other options describe separate CNN components with distinct responsibilities: the loss function (B) is computed at the network's output based on the difference between predictions and ground truth, entirely separate from the kernel's role in feature extraction. Classification (C) is typically performed by fully connected (dense) layers --- often with a softmax activation --- placed after the convolutional feature-extraction stack, not by the kernels themselves. Normalization (D) is handled by dedicated layers such as batch normalization or layer normalization, inserted between convolutional layers to stabilize activations, again a separate mechanism from the convolution operation itself.


Question 9

What is contrastive learning in the context of multimodal deep learning? Pick the 2 correct responses below.

Correct Answer: D. Contrastive learning is a technique used to train deep learning models by comparing similar and dissimilar inputs and optimizing the model to maximize the similarity between representations of similar inputs and minimize the similarity between representations of dissimilar inputs.; E. In a multimodal context, usually, contrastive learning increases the similarity of representations across modalities for the same objects and decreases the similarity of representations across modalities for different objects.
Explanation:

Option D captures the general, task-agnostic definition of contrastive learning: given pairs of inputs labeled as similar (positive pairs) or dissimilar (negative pairs), the training objective pulls positive pairs' representations closer together in embedding space while pushing negative pairs' representations further apart --- typically implemented via losses like InfoNCE, triplet loss, or contrastive loss with a margin. This is the mechanism underlying self-supervised representation learning broadly, not only in multimodal settings.

Option E correctly applies this general principle to the multimodal case: for the *same* object described across modalities (e.g., an image of a dog and the caption 'a dog'), the model should increase representational similarity, since they refer to the same underlying entity; for *different* objects across modalities (an image of a dog paired with the caption 'a cat'), the model should decrease similarity. This is exactly CLIP's training objective, tested elsewhere in this set --- matching image-text pairs pulled together, mismatched pairs pushed apart.

Options B and C both invert this relationship --- B increases similarity for *different* objects and decreases it for *same* objects, and C similarly reverses the correct direction --- describing the opposite of what contrastive learning is designed to achieve, making both clearly incorrect distractors that test careful reading of directionality. Option A is too vague and mischaracterizes contrastive learning as a generative/manipulation technique rather than a representation-learning objective.


Question 10

In LLM evaluation, what does ''zero-shot learning'' refer to?

Correct Answer: D. The model's ability to perform tasks it has not been explicitly trained on
Explanation:

Zero-shot learning describes a model's capacity to correctly perform a task it was never explicitly trained or fine-tuned on, relying instead on knowledge and generalization ability acquired during broader pretraining. For LLMs, this typically means the model is given only a natural-language instruction or prompt describing the task --- with no task-specific labeled examples provided in the prompt at all --- and is expected to produce a reasonable response by generalizing from its pretraining. This is directly analogous to CLIP's zero-shot image classification (covered elsewhere in this set): a model trained broadly can be applied to a new, specific task purely through how the task is described to it, without additional task-specific training.

Option A is a subtly incorrect paraphrase: zero-shot learning is not about the model 'learning from zero examples' during a training process --- it's about applying a model that was never trained for the specific task at all, at inference time. The model isn't learning in the zero-shot moment; it's generalizing from prior training. Option B misapplies 'zero' to training time rather than to task-specific examples --- an unrelated concept. Option C directly contradicts the definition; zero-shot specifically refers to performance *without* task-specific training, not performance *after* extensive training on that task.

Zero-shot is typically contrasted with few-shot learning, where a small number of task-specific examples are included in the prompt to guide the model's response without updating its weights.