ExamVeda
Login
Home
21
The average squared difference between classifier predicted output and actual output.
Discuss
Answer & Solution
Answer: Option A
Solution:
The measure described, which represents the average squared difference between the predicted output of a classifier and the actual output, is known as Option A: mean squared error. Mean squared error is a common metric used to evaluate the performance of machine learning models, with lower values indicating better predictive accuracy.

Option B: root mean squared error is a closely related metric that represents the square root of the mean squared error. It is also used for assessing model performance, and it provides a measure in the same units as the original data.

Option C: mean absolute error measures the average absolute difference between predicted and actual values, but it does not square the differences as in mean squared error.

Option D: mean relative error is not a standard metric for measuring prediction accuracy and is not commonly used in the context of machine learning.

In conclusion, the correct term for the described measure is Option A: mean squared error.
22
Which of the following methods do we use to find the best fit line for data in Linear Regression?
Discuss
Answer & Solution
Answer: Option A
Solution:
In Linear Regression, the method used to find the best fit line for data is Option A: Least Square Error. This technique minimizes the sum of the squared differences (errors) between the predicted values and the actual values in the dataset. The goal is to find the line that minimizes the overall error, making it the "best fit" line for the data.

Option B: Maximum Likelihood is not the primary method for finding the best fit line in Linear Regression. Maximum Likelihood is a statistical method used in other modeling techniques, but it is not the standard approach in Linear Regression.

Option C: Logarithmic Loss is typically associated with logistic regression and classification problems, not Linear Regression.

Option D: Both A and B suggests that both Least Square Error and Maximum Likelihood are used to find the best fit line in Linear Regression. While Maximum Likelihood can be applied in some cases, the primary and standard method is Least Square Error.

Therefore, the correct method for finding the best fit line in Linear Regression is Option A: Least Square Error.
23
Following are the descriptive models
Discuss
Answer & Solution
Answer: Option D
Solution:
Descriptive models in machine learning are used to describe and summarize the relationships and patterns within data without making predictions. Clustering is a descriptive modeling technique that groups similar data points together based on certain characteristics or features.
Association rule is another descriptive modeling technique used to discover interesting relationships between variables in large datasets.
Therefore, Option D is correct because it includes both clustering (Option A) and association rule (Option C) as descriptive models.
24
Assume that you are given a data set and a neural network model trained on the data set. You are asked to build a decision tree model with the sole purpose of understanding/interpreting the built neural network model. In such a scenario, which among the following measures would you concentrate most on optimising?
Discuss
Answer & Solution
Answer: Option C
Solution:
In this scenario, the primary goal is to understand or interpret the neural network model using a decision tree. Fidelity measures how faithfully the decision tree model represents the behavior of the neural network model. It calculates the fraction of instances on which both models provide the same output. Therefore, optimizing Option C ensures that the decision tree model accurately reflects the predictions of the neural network model, aiding in its interpretation.
25
What are common feature selection methods in regression task?
Discuss
Answer & Solution
Answer: Option C
Solution:
In regression tasks, common feature selection methods include:
Option A: correlation coefficient - This method evaluates the strength and direction of the linear relationship between each feature and the target variable.
Option B: greedy algorithms - Greedy algorithms iteratively select features based on certain criteria, such as maximizing predictive performance or minimizing error.
Therefore, Option C encompasses both of these common feature selection methods, making it the correct choice.
26
Regarding bias and variance, which of the following statements are true? (Here 'high' and 'low' are relative to the ideal model.
i. Models which overfit are more likely to have high bias
ii. Models which overfit are more likely to have low bias
iii. Models which overfit are more likely to have high variance
iv. Models which overfit are more likely to have low variance
Discuss
Answer & Solution
Answer: Option C
Solution:
i. False. Models that overfit tend to have low bias because they capture the training data's noise and details, leading to a smaller bias towards the training set.
ii. False. As mentioned in (i), models that overfit typically have low bias, not high bias.
iii. True. Overfitting often leads to high variance because the model captures noise in the training data, resulting in a model that performs well on the training set but poorly on unseen data.
iv. True. Overfitting is characterized by capturing noise and spurious patterns from the training data, leading to low variance since the model's predictions are tightly fitted to the training data.
Therefore, Option C is correct as it includes the true statements iii and iv.
27
Which of the following can only be used when training data are linearlyseparable?
Discuss
Answer & Solution
Answer: Option A
Solution:
Linear hard-margin SVM is designed to find a hyperplane that separates classes with a clear margin and assumes that the data is linearly separable. It aims to maximize the margin while maintaining no training errors, which is only feasible when the data is linearly separable.
Options B, C, and D are not constrained by linear separability:
Option B: linear logistic regression - Logistic regression can be used with both linearly separable and non-linearly separable data.
Option C: linear soft-margin SVM - Soft-margin SVM allows for some misclassifications (errors) and is suitable for non-linearly separable data.
Option D: the centroid method - The centroid method is not specifically designed for linear separability and can be applied to various types of data.
Therefore, the only option that requires linear separability is Option A: linear hard-margin SVM.
28
Wrapper methods are hyper-parameter selection methods that
Discuss
Answer & Solution
Answer: Option C
Solution:
Wrapper methods are hyper-parameter selection methods that involve training multiple models with different subsets of features and selecting the best subset based on performance metrics. They are particularly useful when the underlying learning algorithms are "black boxes," meaning their internal workings are not easily interpretable or understood.
Option A is incorrect because wrapper methods can be computationally intensive since they involve training multiple models.
Option B is incorrect because whether wrapper methods are prone to overfitting depends on their implementation and the data.
Option D is incorrect because wrapper methods can be valuable in certain scenarios, especially when interpretability is not a primary concern.
Therefore, the correct answer is Option C: are useful mainly when the learning machines are "black boxes".
29
Given that we can select the same feature multiple times during the recursive partitioning of the input space, is it always possible to achieve 100% accuracy on the training data (given that we allow for trees to grow to their maximum size) when building decision trees?
Discuss
Answer & Solution
Answer: Option B
Solution:
In decision tree algorithms like ID3 or C4.5, where features can be selected multiple times during recursive partitioning, it is indeed possible to achieve 100% accuracy on the training data. This occurs when the decision tree model perfectly memorizes all the training data, leading to each instance being correctly classified during training. However, achieving 100% accuracy on training data does not necessarily mean the model will generalize well to unseen data. Overfitting can occur, where the model captures noise or anomalies specific to the training data, leading to poor performance on new data.
Therefore, the correct answer is Option A: Yes.
30
In many classification problems, the target dataset is made up of categorical labels which cannot immediately be processed by any algorithm. An encoding is needed and scikit-learn offers at least . . . . . . . . valid options
Discuss
Answer & Solution
Answer: Option B
Solution:
In many classification problems, the target dataset consists of categorical labels that cannot be directly processed by machine learning algorithms. Therefore, encoding techniques are needed to convert categorical labels into a numerical format that algorithms can handle. Scikit-learn offers at least two valid options for encoding categorical labels: Label Encoding and One-Hot Encoding.
Therefore, the correct answer is Option B: 2.