Browse all practice questions for the Casualty Actuarial Society MAS-1 Practice Exam. Search by topic, open any question and review its full explanation, then test yourself in the practice quiz.

Casualty Actuarial Society MAS-1 Complete Practice Test 2026 course image
More practice questions

These questions are part of the practice quiz. Start practicing

  • Which term describes a model that achieves adequate predictive performance using the fewest explanatory variables?
  • If there is a correlation among the error terms, then the estimated standard errors will tend to underestimate the true standard errors.
  • X ~ Exp(theta). any loss over 10,000 will result in a claim payment of only 10,000 due to policy limits. you observe 4 claim payments: 1000, 3100, 7500, 10000. how would you calculate L(theta)?
  • In the context of Poisson residuals, the Pearson residual is defined using which denominator?
  • In general, LOOCV requires fitting a model for a total of n times.
  • The number of events that occur in disjoint time intervals must be independent is a property of counting processes.
  • If the stationary probability of state 0 in a Markov chain is 0.2, what is the expected return time to state 0?
  • Ridge regression tends to shrink coefficient estimates toward zero and typically does not set any coefficients exactly to zero.
  • In Lasso regression, as lambda increases, what happens to the variance of the predictions?
  • If an explanatory variable is uncorrelated with all other explanatory variables, the corresponding VIF would equal what?
  • True or false: as lambda increases from 0 to infinity, the effective degrees of freedom decrease from n to 2.
  • In kernel density estimation, increasing the bandwidth reduces variance but can increase bias; the statement about smoother pdf holds true when bandwidth is larger.
  • Main objective of ridge regression?
  • In ridge regression, what happens to variance as the budget parameter s increases?
  • What is the matrix S defined as in the method for transient absorption probabilities?
  • PCR is useful for performing feature selection.
  • Is the density f(y; θ) = θ y for y > θ a member of the exponential family?
  • For a binary response variable with a continuous explanatory variable, logistic regression is inappropriate.
  • Ridge regression objective can be formulated as minimizing SSR subject to which constraint?
  • In least squares regression, the LOOCV estimate for the test MSE can be calculated by fitting a model once.
  • The deviance is defined as a measure of distance between saturated and fitted model.
  • In ridge regression, what happens to squared bias as s increases?
  • The first principal component direction of the data is the axis along which the observations vary the most.
  • Among the following kernel options for kernel density estimation, which are symmetric?
  • In the Tweedie family, p in (1,2) corresponds to which distribution?
  • In the given model, E[X] equals E[θ].
  • Which statement about k-fold cross-validation is true?
  • Backward stepwise selection cannot be performed on the dataset if n<p.
  • What is the formula for pseudo R-squared?
  • Dimension reduction reduces the number of predictors by projecting onto a lower-dimensional space.
  • In PCA, the first principal component is the direction along which the data vary the most.
  • What distribution do you use when calculating a confidence interval for beta coefficients?
  • Which distribution uses the canonical link inverse squared?
  • In the Tweedie family, p = 3 corresponds to which distribution?
  • In exponential family distributions, canonical form implies a(y) equals y.
  • A parallel system functions as long as one of the components functions.
  • Some regularization methods can also perform variable selection by estimating coefficients to be precisely zero.
  • Greedy Algorithm A is commonly used for which class of optimization problems?
  • In simple linear regression, does the choice of explanatory variable x affect the total sum of squares?
  • If X ~ Exp(λx) and Y ~ Exp(λy) are independent, what is E[min(X,Y)]?
  • True or false: A large value of Mallows Cp indicates a model with a high test error.
  • GAMs allow for non-linear relationships between each predictor variable and the response.
  • In the gambler's ruin scenario, what is Ben's expected final wealth when the total is 75 and the win probability is 0.5?
  • For a sample from an inverse Gaussian distribution, which expression is the MVUE of the mean parameter?
  • The training MSE decreases as model flexibility increases.
  • Given Y claims are made by time t, the unordered times of past events follow what distributions?
  • The cumulative proportion of variance explained increases as more PCs are added.
  • When using the quasi-likelihood approach, is the variance-covariance matrix scaled by an extra dispersion parameter?
  • LOOCV corresponds to k-fold cross-validation with k equal to n.
  • Ridge regression coefficients are not scale equivariant.
  • PCR assumes that the directions in which features show the most variation are the directions that are associated with the target.
  • What is the canonical link for the inverse Gaussian distribution?
  • Lasso regression is able to perform variable selection by forcing some coefficients to be exactly zero.
  • Power parameter in the Tweedie family that corresponds to an inverse-Gaussian distribution.
  • LOOCV uses n training/testing splits, one for each observation left out.
  • Increasing the significance level would decrease the power of a test.
  • Lasso regression performs variable selection by shrinking some coefficients exactly to zero.
  • Best subset selection requires fitting all possible subset models, a total of 2^p models.
  • The smallest possible value of a leverage is 0.
  • In actuarial notation, the symbol a_double_dot_40 denotes:
  • True or false: The critical region of a hypothesis test is determined by the significance level and not by the sample observations.
  • PLS is a dimension reduction method.
  • PCR can reduce overfitting.
  • In a lifetime model where lifetimes are i.i.d. exponential with mean θ, the expected value of the k-th order statistic X_(k) is equal to which expression?
  • Given that the corresponding coefficient estimates for all models with two explanatory variables are the same, which link function will produce a prediction for an observation that is always the greatest?
  • Which kernel density estimator distributes mass uniformly in the neighborhood?
  • Which statement about lasso regression is false?
  • Can a plot of Information be used to visually approximate the MLE of theta?
  • Power parameter in the Tweedie family that corresponds to a Poisson distribution.
  • Relative to the least squares estimates, shrinkage has the effect of reducing bias.
  • In the Tweedie family, p = 1 corresponds to which distribution?
  • Best subset selection results in a nested set of best models with different numbers of predictors.
  • True or false: A small value of span s results in a global fit for a local regression.
  • For an unbiased estimator, the MSE is always equal to the variance.
  • Which expression expresses E[Y] for a distribution in the exponential family?
  • Which type of qualitative variable has categories without a meaningful order?
  • In life-table notation, p_x denotes the probability of surviving from age x to age x+1.
  • What is the primary purpose of fitting a saturated model in generalized linear models?
  • If Ti is the time of the ith event, Pr(T2 > 3) represents what probability?
  • Which statement about the Neyman-Pearson lemma for simple hypotheses is correct?
  • Regression through the origin occurs when the intercept term in the linear equation linking the explanatory variables to the dependent variable is zero or is left out of the equation. Which of the following is true?
  • PCA selects low-dimensional linear surfaces to maximize the captured variance.
  • It is possible to estimate test error by adjusting training error to account for bias due to overfitting.
  • PCA finds a low dimension representation of a dataset that contains as much variation as possible.
  • For a sample from an inverse Gaussian distribution, the MVUE of the mean parameter is:
  • How is multicollinearity detected using the variance inflation factor (VIF)?
  • A predictor uncorrelated with others has VIF equal to 1.
  • Does implementing the quasi-likelihood approach to a GLM change the coefficient estimates?
  • What is the expected value of the top q% of losses?
  • How many of the modeling techniques perform dimension reduction: lasso, PLS, PCA, ridge?
  • In the Poisson context, which statement describes overdispersion?
  • K-fold validation has variance reduction compared to LOOCV.
  • When forming a confidence interval for a proportion, which statistic is used?
  • For a negative binomial distribution with parameter r, the MVUE of the scale parameter is:
  • When forming a confidence interval for a mean with unknown variance, which statistic is used in the calculation?
  • In a fitted values vs residuals graph, heteroscedasticity is indicated by which pattern?
  • Power parameter in the Tweedie family that corresponds to a compound Poisson-Gamma distribution (1<p<2).
  • Which statistic is used to construct a confidence interval for a single population variance?
  • In kernel density estimation, what does the pdf indicate we are looking for?
  • In the context of regression, the statement 'the sample correlation between x and y is equal to the coefficient of determination' is
  • Explanatory variables for Poisson regression besides exposure can be either continuous or categorical.
  • A minimal path set is a minimal set of components whose functioning guarantees the functioning of the system.
  • Using an alternative fitting procedure will likely result in a simpler model.
  • For two independent exponential random variables X and Y, what is E[X | X < Y]?
  • Increasing the significance level would increase the probability of a Type I error.
  • If the number of claims follows a Poisson distribution with rate lambda, what distribution describes the waiting time until the first claim?
  • In the formula f_{2,4} = (s_{2,4} - δ_{2,4}) / s_{4,4} used to compute a probability, what does δ represent?
  • Can a deviance plot be used to visually approximate the MLE of theta?
  • As lambda increases towards infinity, the ridge penalty term has no effect and the estimates become unconstrained.
  • In the Tweedie family, p = 0 corresponds to which distribution?
  • Natural splines, regression splines, smoothing splines, local regression, polynomial regression, and step functions are all types of models that can be used as building blocks for GAMs.
  • In the Tweedie family, p = 2 corresponds to which distribution?
  • What is the MVUE of a binomial distribution?
  • Which statement about bootstrapping and cross-validation is false?
  • Using quasi-likelihood, what must be assumed about the relationship between the mean and the variance?
  • In a likelihood ratio test, which of the following is a correct statement about the null hypothesis?
  • If the hazard rate function decreases with x, the distribution has a heavy tail.
  • PCA provides low-dimensional linear surfaces that are closest to the observations.
  • Residual sum of squares is monotonic with respect to the number of predictors.
  • In a scenario where k-out-of-n with k=1 and n=5, the expected lifetime equals theta times the harmonic sum 1 + 1/2 + 1/3 + 1/4 + 1/5. If theta=2, which value is correct?
  • Which of the following describes the characteristics of a non-homogeneous Poisson random variable?
  • A Markov chain that is irreducible, positive recurrent, and aperiodic is called
  • Which of the following is generally considered unsupervised learning?
  • Ridge regression cannot set any coefficient exactly to zero.
  • Which statement best describes the difference between homogeneous and non-homogeneous Poisson processes?
  • k-fold cross-validation has higher variance than LOOCV when k<n.
  • In life contingencies notation, Ax denotes the present value of what kind of benefit?
  • Evaluating the correlation matrix of predictor variables is a reliable method to detect collinearity.
  • An estimator is consistent whenever the variance of the estimator approaches zero as the sample size goes to infinity.
  • Using an alternative fitting procedure will likely improve prediction accuracy.
  • Performing k-fold cross validation requires fitting a model for a total of k times.
  • If an explanatory variable is uncorrelated with all other explanatory variables, the corresponding variance inflation factor would be 1.
  • A consistent estimator is also unbiased.
  • In the kernel density estimator context, the contribution of a single kernel centered at x_i to the pdf is proportional to which expression?
  • R^2 equal to 1 indicates overfitting.
  • PLS identifies new features in a supervised way by relating them to the target variable.
  • In a Poisson process, the waiting time until the first event is exponentially distributed with a parameter theta equal to what value in terms of lambda?
  • In a Markov chain, a state that cannot be left once entered is called an absorbing state.
  • The incremental variance explained by adding another principal component decreases as more components are included.
  • Power parameter in the Tweedie family that corresponds to a gamma/exponential distribution.
  • In simple linear regression, the F-statistic for the model equals the square of the t-statistic for the slope parameter.
  • Which expression gives the estimated variance of beta_hat in OLS regression?
  • Under X|θ ~ N(θ, σ^2) and θ ~ N(μ, τ^2), the marginal distribution of X is Normal with mean μ and variance σ^2 + τ^2.
  • The first PC is the line in p-dimensional space that is closest to the observations.
  • For modeling hourly bike-sharing usage by day of week, which distribution and link function are most appropriate?
  • Regularized regression methods include ridge and lasso, and their purpose is to prevent overfitting.
  • The fewer positive raw moments that exist, the greater the tail weight.
  • A high leverage point is an outlier.
  • What is the canonical link for the normal distribution?
  • A uniformly MVUE is an estimator such that no other estimator has a smaller variance.
  • In smoothing spline models, increasing lambda affects the bias-variance tradeoff. Which statement is correct?
  • In Greedy Algorithm A, after selecting the lowest-cost assignment, what is done next?
  • When forming a confidence interval for the difference between two means with paired observations, which statistic is used?
  • Decreasing the significance level would increase the probability of a Type II error.
  • Regarding a simple linear relationship, if the irreducible error is zero (e = 0), the 95% confidence interval is equal to the 95% prediction interval.
  • In ridge regression, the shrinkage penalty is applied to all coefficient estimates except for the intercept.
  • What is a good indication of a probability generating function (PGF)?
  • Cayley’s formula gives the number of labeled trees on n vertices. What is that count?
  • True or false: larger values of lambda result in greater effective degrees of freedom for the model.
  • In Lasso regression, as lambda increases, the squared bias of the parameters in the model tends to
  • In a parallel system, the system fails only when all components fail.
  • True or false: a small deviance indicates a poor fit.
  • Which Poisson process type has stationary increments?
  • Which of the following best describes a typical effect of the L1 penalty in lasso regression?
  • What is the MVUE of a normal distribution with a defined variance?
  • What is the canonical link for the Bernoulli distribution in GLMs?
  • True or false: Local regression should not be used in a high-dimensional setting.
  • A saturated model has a deviance of zero.
  • Greedy Algorithm B uses which sequence of k values?
  • The sum of the leverages across all observations must equal the number of explanatory variables.
  • An annuity-due issued to a 40-year-old that pays 1 each year until death or age 60, whichever comes first, has an actuarial present value given by which expression?
  • Backward stepwise selection cannot be used when the dataset has as many predictors as observations (n < p+1).
  • With least squares regression on a dataset with n observations, the LOOCV estimate for the test MSE can be calculated by fitting a model once.
  • Compared with lasso regression, ridge regression is generally harder to interpret because it retains all predictors in the model.
  • Which Excel expression yields the p-value for a chi-square test?
  • Before applying ridge regression, predictors should be standardized because ridge is not scale invariant.
  • Which of the following is NOT listed as a potential cause of unreliable mean squared error estimates?
  • Removing any component from a minimal path set will break the guarantee.
  • The Tweedie distribution is particularly useful when data include zeros and continuous positive values that can be viewed as which distribution?
  • Mallows Cp is an unbiased estimate of the test MSE if calculated using an unbiased estimate of variance.
  • LOOCV bias is lower than k-fold CV bias.
  • Removing high leverage points in a linear regression model primarily affects which aspect of the fitted model?
  • A consistent estimator cannot be biased.
  • The Excel function CHISQ.DIST.RT can be used to compute the p-value for a chi-square statistic.
  • Which property ensures a Markov chain has a unique stationary distribution and convergence from any starting state?
  • As more variables are added to a linear model, the training MSE generally does what?
  • If a simple linear model for a probability predicts values outside the [0,1] interval, what is the standard corrective approach?
  • When selecting the optimal model, which criterion is preferred?
  • What is the typical statement about removing outliers on model fit?
  • If both classes were transient, after some time, the chain would not be in either class.
  • If two statistics have the same mean, the one with smaller variance is called the efficient estimator.
  • The estimated variance of a coefficient in OLS is given by which expression?
  • Which distribution has a constant hazard rate?
  • A gamma random variable with alpha = 2 and theta = 1 to be when simulating random variables?
  • The deviance is useful for testing the significance of explanatory variables in nested models.
  • What is the canonical link for the gamma/exponential distribution?
  • A 5-state Markov chain with two classes {0,1,3} and {2,4} is ergodic.
  • Power parameter in the Tweedie family that corresponds to a normal distribution.
  • Overdispersion occurs when the observed variance is larger than the mean for a Poisson model, or when the observed variance is larger than the calculated variance for a binomial model. Which statement best captures overdispersion in common models?
  • True or false: Mallows Cp and AIC are proportional to each other.
  • True or False: R-squared is a good measure for model comparison.
  • What is the likelihood ratio critical region for testing H0 against H1?
  • Mallows Cp and AIC are proportional to each other; in fact, they are equal.
  • In a PCA performed on a data set with 50 observations and 3 independent continuous variables, which statement is true?
  • The variance of error terms doesn't have to be constant.
  • In Lasso regression, as the regularization parameter lambda increases, what happens to the number of predictors selected?
  • What's the formula for Cov(X,Y)?
  • In a smoothing spline model fit to data, what happens to bias as the tuning parameter lambda increases?
  • Using an alternative fitting procedure makes results easier to interpret.
  • Which sequence of steps correctly calculates the probability of moving from transient state 2 to state 4 in a Markov chain?
  • True or false: The model containing all predictors will always have the smallest residual sum of squares and largest R-squared.
  • The MVUE of the scale parameter beta in Gamma(shape=alpha, scale=beta) is which expression?
  • Which statement best indicates a variable is statistically significant in a standard hypothesis test?
  • Cross-validation is used to measure the accuracy of a parameter estimate.
  • A GLM with the same distribution and link function as the model of interest that has the max number of parameters that can be estimated; assists in assessing model adequacy is called what?
  • In a local regression model, increasing the span s will typically produce what effect on the fitted curve?
  • A uniformly MVUE is defined as an estimator such that no other estimator has a smaller variance.
  • what is the excel equation for solving for the cdf in gaussian kernel estimation?
  • Which formula correctly expresses Var(X+Y) accounting for dependence?
  • Dimension reduction involves projecting the p predictors into an M-dimensional subspace by computing M distinct linear combinations of the variables and utilizing them as predictors for fitting a linear regression model.
  • In nested models, the deviance is useful for testing the significance of explanatory variables.
  • In a hierarchical normal model where X|θ ~ N(θ, 100^2) and θ ~ N(800, 50^2), what is E[X]?
  • The pseudo R-squared is computed as one minus the ratio of the model log-likelihood to the null log-likelihood.
  • If n = 4, how many minimal cut sets are there?
  • True or false: all states in an irreducible Markov chain are recurrent.
  • Is a regression tree an example of supervised or unsupervised learning?
  • In smoothing spline models, increasing lambda has what effect on the bias-variance tradeoff?
  • How many minimal cut sets are there for a random graph with n nodes?
  • What's the formula for Var(X+Y)?
  • A counting process possesses independent increments if the number of events between s and t is independent of the number between t and t+u for all u>0.
  • Pr(X>x) with a conditional distribution? (Law of Total Probability)
  • Which of the following best describes KNN?
  • A Markov chain with two communicating classes is not irreducible.
  • In a linear regression model, the leverage for each observation is guaranteed to lie between 1/n and 1.
  • Which of the following statements best describes a non-homogeneous Poisson process?
  • Collinearity can exist among three variables even if no single pair shows a high correlation.
  • Using the ILT to price life insurance policies, the lower bound for the number of deaths during a period is given by which expression?
  • Rank the following tools by flexibility in descending order: spline, linear regression, ridge regression.
  • Greedy Algorithm B for optimization problems is described as which approach?
  • K-fold validation has a computational advantage over LOOCV when k < n.
  • In ridge regression, which parameter is not subject to shrinkage?
  • Shrinkage reduces variance at the cost of a small increase in bias.
  • K-fold validation has an advantage over LOOCV in variance reduction.
  • For a k-out-of-n system with iid exponential components, the expected lifetime is theta * sum_{i=k}^n 1/i. If theta=2, n=5, k=3, what is the expected lifetime?
  • Ordinary least squares estimators are inherently unbiased.
  • N(t) must be greater than or equal to 0 is a property of counting processes.
  • Is boosting an example of supervised or unsupervised learning?
  • Which norm is used in the penalty term for lasso regression?
  • Which inequality defines the likelihood ratio test critical region?
  • In testing whether a source is significant, the test statistic is the mean square of that source divided by the MSE of the model that has the most predictors.
  • Adjusted R^2 equal to 1 indicates overfitting.
  • In the exponential family, the expression for E[Y] can be written as E[Y] = - c'(θ) / b'(θ).
  • In local regression, increasing the span parameter makes the fit more global rather than local.
  • A series system functions only when all components function.
  • Which step is typically performed first when computing expected sojourn times for a Markov chain with transient states?
  • In ridge regression, the sum of squares of beta is bounded above by s. Which statement best captures this constraint?
  • Var(S) for S = sum_{i=1}^N X_i with N ~ Poisson(lambda) and i.i.d. X_i is equal to lambda * E[X^2].
  • When applying the quasi-likelihood approach to a GLM, does it change the point estimates of the coefficients?
  • In generalized linear models, the statement 'the saturated model has the highest possible deviance' is true or false?
  • Increasing model flexibility decreases variance.
  • A logit model applies when the explanatory variables are both continuous and categorical.
  • In a regression with p predictors and an intercept, the trace of the hat matrix (sum of leverages) equals:
  • Homoscedasticity occurs when what condition holds?
  • In simple and multiple linear regression, the maximum likelihood estimator for the regression coefficients coincides with the ordinary least squares estimates when residuals are normally distributed.
  • Deviance can be used to test the significance of explanatory variables in nested models.
  • To form a confidence interval for the ratio of variances between two populations, which statistic is used?
  • Which statement about deviance is correct?
  • In Tweedie distributions, data are appropriate when data include zeros and continuous positive values that can be viewed as which distribution?
  • GAMs are a useful representation if we are interested in inference, since you can examine the effect of the predictor variables on the response while holding all of the other predictor variables constant.
  • Which statement correctly differentiates homogeneous and non-homogeneous Poisson processes regarding stationary increments?
  • A biased estimator can be consistent.
  • LOOCV requires fitting a model a total of n times.
  • Residual plots are a useful graphical tool for identifying non-linearity.
  • When forming a confidence interval for a population variance, what kind of statistic is used?
  • LOOCV tends to overestimate the test error rate in comparison to validation set approach.
  • Are all Poisson processes characterized by stationary and independent increments?
  • In ridge regression, coefficients shrink toward zero but are typically not exactly zero.
  • The deviance is defined as a measure of distance between saturated and fitted model.
  • Which statement about training set MSE versus test MSE is true?
  • PLS is a subset selection method.
  • An unbiased estimator is considered a consistent estimator if the variance of the estimator converges to 0 as n approaches infinity.
  • Ridge regression coefficients are not scale equivariant.
  • Which statement is true about backward stepwise selection?
  • What is the form of the likelihood function for two independent populations with different density parameters?
  • When a lognormal X ~ Lognormal(mu, sigma^2) is scaled by a positive constant c, which distribution describes cX?
  • What is the formula for a Pearson residual?
  • With E(N)=150, Var(N)=100, z=1.96, what is the ILT lower bound for the number of deaths?
  • For an exponential distribution with complete data, what is the maximum likelihood estimator of the mean parameter?
  • Given the minimal path sets {1,2,5}, {1,3,4}, {2,3,5}, {3,4,5}, which of the following is a minimal cut set?
  • The maximum possible value of the standard error of the population variance is achieved under which condition?
  • What penalty term does ridge regression use?
  • Forward stepwise selection requires fitting 1 + (1/2)(p*(p+1)) models.
  • In ridge regression, training error as s increases?
  • In ridge regression, test error as s increases?
  • A logit transformation helps in reducing heteroscedasticity.
  • R^2 is the fraction of variation in y about the mean of y that's explained by the linear relationship with x.
  • True or false about a smoothing spline model fit to data using the tuning parameter lambda: larger values of lambda result in smoother splines.
  • The smoothness of a continuous predictor variable in a GAM can be summarized by degrees of freedom.
  • LOOCV requires fitting a model a total of n times.
  • In a Galton-Watson branching process with offspring probabilities P_j, the extinction probability π0 satisfies which equation?
  • If theta-hat is unbiased and efficient, then theta-hat is the MVUE.
  • For an exponential random variable X, what is E[X | X > a]?
  • For a Negative Binomial distribution with parameter r, which expression is the MVUE of the scale parameter?
  • When deciding between a regression spline and local regression, which component must be considered for a regression spline but not for local regression?
  • What is the test statistic for testing the significance of a single parameter in a regression model?
  • What happens to the variance-covariance matrix when implementing the quasi-likelihood method?
  • What is the sum of the leverages across all observations in a linear regression with an intercept and p predictors?
  • Which statement about LOESS span and smoothing is true?
  • Which statement best describes how lasso influences model sparsity?
  • Which expression is the MLE for the exponential distribution when data are censored and truncated?
  • How many minimal path sets are there for a random graph with n nodes, according to the given material?
  • Using all possible PCs provides the best understanding of the data.
  • What happens to training mean squared error as model flexibility increases?
  • In a one-dimensional symmetric random walk, where the probability of moving in either direction is 0.5, are all states recurrent?
  • What is a commonly used model for times to failure (or survival times)?
  • If X_(k) is the kth order statistic from an iid sample from Uniform(0, θ), which distribution does X_(k) follow?
  • It is possible to directly estimate a model's test error using a validation set or cross-validation.
  • When forming a confidence interval for a mean with a known variance, which statistic is used in the calculation?
  • In a three-state Markov chain with states 0, 1, 2 and starting in 0, what is the formula for the expected number of steps to return to state 0?
  • When forming a confidence interval for the difference between two means with known variances, which statistic is used?
  • To model a non-negative response with an unbiased estimate, which error structure and link function combination is most appropriate?
  • For a Bernoulli response in a generalized linear model, which set of link functions can be used?
  • If all regression errors are identically zero, what is the R-squared value?
  • After standardizing the predictors, which statement about the PLS first direction is correct?
  • In ridge regression, increasing the tuning parameter lambda strengthens the penalty and shrinks coefficients toward zero.
  • Kernel density estimation is used to estimate which component of a distribution?
  • The efficiency of theta-hat is the estimator's variance divided by the Rao-Cramer lower bound.
  • Adjusted R-squared equal to 1 indicates overfitting.
  • Which statement about the relationship between Mallows Cp and AIC is most consistent with the material?
  • What does the likelihood ratio test (LRT) test?
  • In PCA, the first few principal components are often sufficient to get a good understanding of the data.
  • If n = 4, how many minimal path sets are there?
  • In ridge regression, which statement is true?
  • When comparing two means assuming known variances for both populations, which statistic is used to form the confidence interval?
  • What is the canonical link function for Poisson regression?
  • The deviance for normal distributions is proportional to the residual sum of squares.
  • SSE equals zero indicates overfitting.
  • Cluster analysis is typically categorized as which type of learning?
  • Poisson regression models are capable of handling varying exposure by allowing the exposure term to differ across observations.
  • In quasi-likelihood, how is the variance-covariance matrix adjusted?
  • For a natural cubic spline, what is the number of degrees of freedom?
  • A large value of a leverage indicates the presence of an outlier.
  • Is PCA an example of supervised or unsupervised learning?
  • Which statement describes a common effect of high dimensionality on MSE estimates?
  • Which statistic would you use to form a confidence interval for the difference of two means when the variances are unknown?
  • All collinearity problems can be detected by inspection of the correlation matrix.
  • A distribution with support depending on theta cannot be a member of the standard exponential family.
  • According to the Central Limit Theorem, what is the limiting distribution of the sampling distribution of the sample mean?
  • In binomial data, overdispersion manifests as the observed variance exceeding the binomial variance, which is expressed as which formula?
  • To determine asymptotic unbiasedness, which condition must hold?
  • A Poisson process with a constant rate is called what?
  • Which statement about the kth order statistic from Uniform(0, θ) is correct?
  • In smoothing splines, setting lambda to zero corresponds to no penalty for roughness.
  • In regression analysis, does an R-squared value of 0 indicate overfitting?
  • Conditional on θ, X follows a Normal distribution with mean θ and variance 100^2.
  • If a distribution isn't in canonical form, can there be a natural parameter?
  • Which method is used to select the appropriate level of model flexibility?
  • Lasso regression tends to outperform ridge regression in terms of bias, variance, and MSE.
  • True or false: if all states in a finite Markov chain are recurrent, the Markov chain is irreducible.
  • Ridge regression uses an L2 penalty on coefficients.
  • what is the excel equation for solving for the pdf in gaussian kernel estimation?
  • Both ridge and lasso regression are regularized methods.
  • In Poisson regression, which statement about the variance-mean relationship is true?
  • What could be added to a linear regression model to avoid multicollinearity?
  • Which term describes a Markov chain that has only one communicating class?
  • Dimension reduction involves projecting into a lower-dimensional subspace using M linear combinations.
  • In a 10-state Markov chain, which property ensures that all states communicate?
  • Which type of qualitative variable has categories with a meaningful order?
  • Using Cook's distance with a unity threshold, an observation is influential if which condition holds?
  • For theta=0.5, n=4, k=2, what is the numeric value of E when E = theta * sum_{i=k}^n 1/i?
  • A limitation of GAMs is that interactions cannot be added to the model.
  • If the number of PLS components equals the number of predictors in OLS, the forecasted values from both methods are what?
  • Deviance in generalized linear models and the chi-square distribution: deviance follows a chi-square distribution for all models in the exponential family.
  • Poisson regression models incorporate a logarithmic link function.
  • Only w-1 dummy variables are needed to represent w classes of a categorical predictor.
  • In kernel density estimation using a Gaussian kernel, the width of the neighborhood is infinite.
  • True or False: The F-statistic used to test a predictor in regression is computed as the ratio of its mean square to the mean square error from the model with the most predictors.
  • PCA can be used for data visualization.
  • PLS identifies new features in an unsupervised way by approximating the original predictors, similar to PCA.
  • Which of the listed modeling procedures performs variable selection?
  • Which model uses more parameters for the same number of knots: a linear spline with k knots or a natural cubic spline with k total knots?
  • In ordinary least squares, the variance of each coefficient estimate is given by which expression?
  • R^2 is the ratio of the regression sum of squares to the total sum of squares.
  • How is Greedy Algorithm A described for optimization problems?
  • Deviance is a measure used to assess the quality of fit for nested models.
  • What is the test statistic for testing the equality of two variances?
  • If state 2 is positive recurrent, then state 4 must be positive recurrent.
  • Deviance is minimized to obtain the best-fitting generalized linear model; in general, lower deviance indicates a better fit.
  • Forward stepwise selection cannot be used in high-dimensional settings.
  • In Gaussian kernel density estimation, the data point x_i represents what in the kernel sum?
  • How do you standardize a residual?
  • Poisson regression assumes that the mean equals the variance of the response variable.
  • To determine if a function should be used as a link function for a GLM, check if the function is monotone and differentiable.
  • What is the only continuous distribution listed with a finite support?
  • Which statement about lasso regression compared to ordinary least squares is true?
  • In regression with Gaussian errors, a large value of Mallows' Cp indicates a model with a low test error.
  • The sum of leverages across observations equals p+1.
  • In actuarial notation, A_x denotes the present value of a 1-unit death benefit payable at the end of the year of death for a life aged x.
  • Which statement correctly describes the relationship between MSE and the true parameter?
  • Why do we use w-1 dummy variables for a categorical predictor with w levels?
  • What represents the moment generating function of Y evaluated at t=1, My(1)?
  • The MVUE is defined as the unbiased estimator with the minimum variance.
  • Shrinkage fits a model involving a subset of predictors with the estimated coefficients shrunken towards zero.
  • How is overdispersion detected in a generalized linear model?
  • For a distribution in the exponential family, E[Y] equals negative derivative ratio - c'(θ) / b'(θ).
  • In a series system, the system functions only if every component is functioning.
  • What best describes the type of problem Angela is solving by clustering shoppers to target ads?
  • Does the quasi-likelihood method change the coefficient estimates?
  • ANOVA is a useful approach for analyzing the means of groups of continuous response variables, where the groups are categorical.
  • Which statement about ridge regression is true?
  • Loadings for the first direction are proportional to covariances between the response and each standardized predictor.
  • Poisson regression models assume exposure is constant when modeling the rate at which events occur.
  • In ridge regression, irreducible error as s increases?
  • Which of the following expresses ridge regression as a constrained optimization problem?
  • In linear regression with p predictors and an intercept, the sum of leverages equals which of the following?
  • If X ~ Uniform(m, n), what is the distribution of (X | X > pi_q)?
  • The efficiency of an estimator is defined as the Rao-Cramer lower bound divided by the estimator's variance.
  • How do you identify overdispersion in a model?
  • Regarding kernel density estimation: the larger the bandwidth, the smoother the estimated pdf is.
  • Ridge regression uses an L2 penalty, while lasso regression uses an L1 penalty.
  • In a normal linear model, the scaled deviance is equal to which of the following?
  • Do the beta_hat values of a ridge regression procedure provide unbiased estimators of the corresponding beta model parameters?
  • What is MGF of Y evaluated at t=1?
  • In regression context, if there is no linear relationship between x and y, the scatterplot will typically show a random pattern.
  • In a regression model that includes an intercept, the sum of residuals is:
  • Unlike the validation set approach, the k-fold cross-validation approach uses all observations to train the model.
  • Which interval quantifies the possible range for a future observation Y given X?
  • Is an absorbing state considered transient or recurrent?
  • Ridge regression outperforms lasso when the response is a function of many predictors, all with coefficients of roughly equal size.
  • In marketing analytics, clustering to segment shoppers is best described as which type of learning?
  • What is the MVUE of sigma^2 for a normal distribution with unknown mean and unknown variance?
  • As model flexibility increases, the test MSE monotonically decreases.
  • Residual sum of squares is a suitable metric for selecting the best model among models with different numbers of predictors.
  • In simple linear regression, a random pattern in the scatterplot of y against x indicates that R^2 is near zero.
  • Which methods guarantee a nested sequence of models as predictors are added or removed?
  • Using a canonical link function in a GLM, are the estimates unbiased or biased?
  • In a Poisson regression model, what is the offset term?
  • The cumulative proportion of variance explained cannot decrease when more PCs are added.
  • For a normal distribution, the deviance is proportional to the residual sum of squares.
  • For a sample from Uniform(a,b), the expected minimum (k = 1) is a + (b - a)/(n + 1). Which expression correctly represents this?
  • For the acceptance-rejection method, what's the inequality for f(y)/ (c g(y)) relative to a Uniform(0,1) random variable U?
  • True or false: Mallows Cp is an unbiased estimate of the test MSE if its variance is calculated using an unbiased estimate of the variance.
  • Which sequence leads to the asymptotic variance of θ?
  • In ordinary least squares, the sum of residuals equals which value?
  • In kernel density estimation, after calculating the kernel contributions k_i(x) from each observation, how is the density estimate at x formed?
  • If Y is complete sufficient for theta and g(Y) is unbiased for theta, then g(Y) is the MVUE with the smallest variance.
  • What is Mallows Cp equation?
  • For X ~ Uniform(m,n), the conditional distribution X | X > pi_q is Uniform(pi_q, n) provided pi_q lies in (m,n).
  • What is the probability that an observation is not selected for a bootstrap sample?
  • Collinearity reduces the accuracy of the estimates of regression coefficients and may make it harder to reject the hypothesis that beta_j = 0.
  • Mallows Cp involves SSE, p, and MSEfull.
  • In the transient-state fundamental matrix S = (I - PT)^{-1}, what does the entry S_ij represent?
  • In curtate life expectancy, ex_(curtate) relates to survival probability p_x and the life expectancy at age x+1 by which expression?
  • Which statement about the exponential family and canonical form is true?
  • The logit model is appropriate when the response variable is binary.
  • SSE equal to 0 indicates overfitting.
  • A logit model gives numerical results that are quite similar to those given by the probit model.
  • In the context of the material, when estimating an exponential distribution, the MLE of the mean equals the sample mean.
  • Variance refers to the error arising from the assumptions made in the statistical learning tool.
  • Which description matches the constrained form used in the lasso alternative objective?
  • If Y is a complete sufficient statistic for theta and g(Y) is an unbiased estimator of theta, then g(Y) is the MVUE and has the smallest possible variance among all unbiased estimators.
  • K-fold validation has an advantage over LOOCV in bias reduction.
  • A Markov chain that has a limiting distribution is described as which type?
  • Simon uses a statistical learning method to estimate the number of ears of corn produced per acre. He applies the same method to multiple training data sets and results are similar but not identical. What best describes this method?
  • Which distribution and link function should be used for slices of pizza sold at a convenience store based on distance to the city center?
  • Bias refers to the error arising from the method's sensitivity towards the training data set.
  • Which expression is the MSE of an estimator?
  • In simple linear regression, R^2 equals the square of the sample correlation coefficient r. Which option is true?
  • What makes a good argument for choosing LOOCV over 5-fold CV?
  • Which statement is NOT an assumption when using pooled variances for two-sample tests?
  • What is the expected value of the kth order statistic from a Uniform(a,b) distribution?
  • In the Poisson distribution, which statement is true about the relationship between the mean and the variance?
  • Consistency of an estimator is characterized by the variance converging to zero as the sample size grows.
  • Which method is used when a problem mentions the 'best critical region'?
  • In supervised learning, the variance and the squared bias are inversely related.
  • The Cramer-Rao lower bound for the variance of all unbiased estimators of theta equals:
  • Which of the following best describes the objective of lasso regression?
  • A logit model applies when the response variable counts the number of events occurring.
  • Mean squared error (MSE) is defined as the expected squared difference between the estimator and the true parameter. Which statement is true?
  • Which distribution is commonly used to model life data due to its flexible hazard function?
  • In a Markov chain, a state whose probability of returning is less than 1 is called a
  • True or false: We should choose a model with a low training error when selecting the optimal model.
  • The standard error of regression uses degrees of freedom equal to n-2.
  • For a given dataset, the number of variables in a Lasso regression model will always be greater than or equal to the number of variables in a Ridge regression model.
  • Best subset selection requires fitting all (2 choose p) models for each possible combination of p predictors.
  • With a cubic spline, what must match at the knot?
  • In simple linear regression, which interval estimates E(Y|X)?
  • When forming a confidence interval for the ratio of two variances, what kind of statistic is used?
  • For a cubic spline with one knot, how can we connect the two pieces of the equation to find coefficient estimates?
  • Leave-one-out cross-validation is a special case of k-fold cross-validation where k equals the number of observations.
  • Under a saturated model, the predicted value for a given observation is
  • What is the MVUE of a normal distribution with a defined mean?
  • What is the canonical link for the Poisson distribution?
  • Is a cluster analysis an example of supervised or unsupervised learning?
  • Minimal cut sets must have at least one component from each minimal path set.
  • Under the saturated model, what is the predicted value for each observation?
  • Which statement best describes how model flexibility affects variance and bias?
  • The statement that the proportion of variance explained by an additional principal component increases as more PCs are added is true or false?
  • Which statement about leverages in a linear model with intercept and p explanatory variables is correct?
  • For X ~ Exp(λx) and Y ~ Exp(λy) with 1 < X < Y, what is E[X | 1 < X < Y]?
  • The squared bias increases as the method's flexibility decreases.
  • If theta-hat is the MVUE, theta-hat is efficient.
  • Both lasso and ridge regression shrink coefficients toward zero, but lasso can force some coefficients to be exactly zero.
  • In dummy coding for a categorical predictor with four levels, how many dummy variables are needed if one category is used as the baseline?
  • With 3 original variables, what is the maximum number of principal components that can be extracted?
  • What is the MVUE of a Poisson distribution?
  • Deviance is a useful measure of goodness of fit for all models in the exponential family.
  • Which of the following increases monotonically as model flexibility increases?
  • Which statement best captures the meaning of stationary and independent increments?
  • Which of the following statements is true about one-dimensional and two-dimensional symmetric random walks?
  • In GLMs, the primary consideration for choosing between a Poisson model with a log link and a Gaussian model with an identity link is the distribution of the response variable.
  • In PCA, the third principal component is orthogonal to the first principal component.
  • Can a plot of the score function be used to visually approximate the MLE of theta?
  • The statement 'The logit link corresponds to a logistic distribution, and the probit link corresponds to a standard normal distribution' is true.
  • In kernel density estimation, what does the cdf indicate we are looking for?
  • Which statement about a simple linear relationship is true?
  • In simple linear regression, the least squares line passes through the point (x-bar, y-bar).
  • If two sampling distributions have the same mean, the one with smaller variance is called what?
  • Does the quasi-likelihood approach change the coefficient estimates?
  • True or False: Power is the probability of rejecting the null hypothesis, assuming its false.
  • K-fold cross validation requires fitting a model for a total of k times.
  • Bootstrapping can be used to select the appropriate level of model flexibility.
  • If we want the expected number of time periods a chain is spent in state j given it started in state i, what are the steps for calculating this?
  • For paired observations, which statistic is used to form a confidence interval for the difference in means?
  • In k-fold cross-validation, the model is fitted a total of k times.
  • Which method is used to measure the accuracy of a parameter estimate?
  • Which process is characterized as a counting process with integer-valued counts?
  • What is MSE(estimator)?
  • What is the test statistic for a likelihood ratio test?
  • Overfitting causes the training error to underestimate the test error.
  • If N ~ Poisson(lambda) and S = sum_{i=1}^N X_i where X_i are independent of N with E[X^2] finite, what is Var(S)?
  • Which statement is true about the scale behavior of ridge regression?
  • The Cramer-Rao lower bound for the variance of unbiased estimators is given by which expression?
  • In a linear model with an intercept and p explanatory variables, the leverage for each observation must be between which values?
  • True or false: For the same total number of knots k, the natural cubic spline uses fewer degrees of freedom than the cubic spline.
  • Which statement about ridge regression is true?
  • For Poisson processes, counts in disjoint intervals are independent.
  • In PCR, it is common to use only the first few principal components to predict the response.
  • The tuning parameter for ridge regression can be selected using cross-validation.
  • PCA serves as a tool for data visualization.
  • Ridge regression shrinks coefficients toward zero and can never set any coefficient exactly to zero.
  • True or false: Ridge regression is less flexible than OLS and thus results in an improved prediction accuracy when its increase in squared bias is less than its decrease in variance.
  • In computing the first direction, PLS places the highest weight on the variables that are most strongly related to the response.
  • The irreducible error variance in this model is 100^2.
  • Which ordering of the degrees of freedom used by the three spline models is correct from most to least, given a linear spline with k knots, a cubic spline with k knots, and a natural cubic spline with k total knots (k minus 2 interior knots)?
  • LOOCV is a special case of k-fold cross-validation.
  • N(t) doesn't have to be an integer.
  • How can you identify the mode using a probability density function?
  • True or false: since training error can be a poor estimate of the test error, RSS and R-squared are not suitable for selecting the best model.
  • A chain with only one class is called an irreducible chain.
  • n^(n-2) is associated with the count of which standard combinatorial object?
  • Which statement about Mallows Cp and AIC is supported by the material?
  • What penalty term does lasso regression use?
  • True or false: Local regression is a memory-based procedure.
  • If the branching process starts with n individuals, the extinction probability is π0^n.
  • In simple linear regression, the statement that the sample correlation between x and y equals the coefficient of determination R^2 is
  • In least squares LOOCV, the LOOCV error can be computed from a single fitted model using residuals and leverage.
  • A probit link is a valid alternative to the logistic link for binary outcomes.
  • How do you test for time reversibility of a Markov chain?
  • In a 5-state Markov chain with two classes {0,1,3} and {2,4}, at least one of the two classes must be recurrent.
  • When scaling X ~ Lognormal(mu, sigma^2) by a positive constant c, which distribution describes cX?
  • In the gambler's ruin scenario with total wealth 75, Ben starts with 40 and Allison with 35. Which expression correctly computes the expected final wealth of Ben?
  • Which kernel density estimation fact is true?
  • A biased estimator can only be inconsistent.
  • When forming a confidence interval for the difference between two means with unknown variances, which statistic is used?
  • Which distribution is associated with the canonical link that is inverse?
  • In a Markov chain, positive recurrence is a property that holds for all states in a given communicating class.
  • What does removing outliers do to a linear regression model?
  • What is the formula for the degrees of freedom in a chi-square goodness-of-fit test with g groups and p estimated parameters?
  • Ridge regression shrinks the coefficient estimates, which has the benefit of reducing the bias.
  • When forming a confidence interval for the difference of two proportions, which statistic is used?
  • In Poisson regression, the Pearson residual is computed as (y - mu_hat) / sqrt(mu_hat). What does mu_hat represent in this context?
  • For modeling a binary outcome such as hospitalization, which distribution and link are most appropriate?
  • In the same model, what is Var(X)?
  • Which of the following is generally considered supervised learning?
  • What is the symbol used for the expected number of time periods a chain is in state 2 given the chain starts in state 1?
  • In an exponential family distribution, the sufficient statistic for the natural parameter mu, given observations x1,...,xn, is which of the following?
  • For a Gamma distribution with known shape parameter alpha and i.i.d. samples, what is the MVUE of the scale parameter beta?
  • When should you use pooled variances for confidence intervals (or tests)?
  • In a Poisson process, increments over disjoint time intervals are what property?
  • In a Markov chain, a state that is guaranteed to be revisited eventually is called a
  • In PCR, is it recommended to standardize each predictor prior to generating principal components?
  • The law of total variance states Var(X) = E[Var(X|θ)] + Var[E(X|θ)].
  • When given a CDF of a distribution, how do you obtain the CDF of the kth order statistic Y_(k)?
  • Is KNN an example of supervised or unsupervised learning?
  • In a residuals vs fitted values plot, heteroscedasticity is indicated by which pattern?
  • Power is defined as 1 minus the probability of a Type II error.
  • Which of the following statements about the prediction interval and the range it measures is true?
  • what is the neyman-pearson theorem for hypothesis testing?
  • Which expression represents Cov(X,Y) in terms of expectations?
  • Positive recurrence is a class property; if state 2 is positive recurrent, then state 4 must be positive recurrent.
  • For X ~ Uniform(2,5), conditioning on X > 3 yields X | X > 3 ~ Uniform(3,5).
  • Which statement best describes an ergodic Markov chain?
  • Subset selection is used to identify a subset of the predictors and then fit a model using least squares on the reduced set of variables.
  • In least squares regression, the LOOCV estimate for the test MSE can be calculated by fitting a model once.
  • True or false: Ridge regression is less flexible and thus results in an improved prediction accuracy when its decrease in variance is less than its increase in squared bias.
  • Which model form will both have a discontinuity in the fitted curve and most likely overfit the data when predicting height from shoe size?
  • Cp, AIC, BIC, and adjusted R-squared are used to adjust which type of error when evaluating models for model size?
  • Ordinal variables are a type of continuous explanatory variable.
Subscribe

Get the latest from Examzify

You can unsubscribe at any time. Read our privacy policy