Statistics and Data Science Seminar : Past Events
Past Seminars
The following seminars have already happened, you may instead view upcoming seminars in this series.
Aug. 31, 2005
Professor Dibyen Majumdar :
3 p.m. in SEO 512
Abstract
Professor Sujit Kumar Mitra was a luminary in the world of statistics and linear
algebra. The clarity and brilliance of his results and proofs have been seldom
matched. I will attempt an appreciative reflection on his work through examples
on asymptotics of the contingency chi-square, characterization of the Wishart
distribution, MINQUE theory, matrices and the g-inverse. The talk will be at the
graduate student level.
Sept. 14, 2005
Professor Lan Zhang :
3 p.m. in SEO 512
Abstract
The availability of high frequency data for financial instruments has opened the
possibility of accurately determining volatility in small time periods, such as
one day. Recent work on such estimation indicates that it is necessary to treat
the data with a hidden semimartingale model, typically by the addition of
measurement error. We review the emerging theory on this subject, including two-
and multiscale sampling.
Sept. 21, 2005
Xin Fang :
3 p.m. in SEO 512
Abstract
Both pharmacodynamic (PD) Emax models and pharmacokinetic (PK) one-compartment
models have been well studied and widely used in pharmaceutical industry to
obtain the ADME (absorption, distribution, metabolism and excretion) and
evaluate the efficiency of a drug candidate. However, the implementation of
experimental designs on Emax models is difficult since concentrations of the
drug in blood are uncontrollable. In this talk, two models that combine the PD
and PK models are discussed to address this problem. With the current models,
one can obtain better estimations of the parameters of the ADME and efficiency
of a drug through implementing a locally D-optimal design. A class of robust
designs is also investigated for comparison. Simulation results show that the
robust designs have better statistical properties when sample size is large such
as in clinical trial III, while the locally D-optimal designs are more proper
when sample size is small as in clinical trials I & II.
Sept. 28, 2005
Professor Balakrishna Hosmane :
3 p.m. in SEO 512
Abstract
Evaluation of new drugs for unwanted effects on electrical properties of the
heart is receiving heightened attention from pharmaceutical companies and
regulatory agencies. This attention arises from recent scientific research that
links drug effects in cellular ion channels to changes in electrical
characteristics of the electrocardiogram (ECG) that predict clinically important
cardiac arrhythmias. An assessment of non-inferiority of the higher dose to
placebo is performed in four-period cross-over study as well as four-group
parallel study by the union- intersection test within the framework of a linear
mixed effects analysis. For the purposes of planning such a study, the joint
distribution of the estimate of the difference in means of high dose of
investigational drug and the placebo was derived. The power of thorough QT/QTc
study evaluated using the joint distribution and the simulation study were quite
close.
Oct. 5, 2005
Payal Pagni :
3:30 p.m. in SEO 512
Abstract
We will discuss a two-phase sampling situation where the first phase consists of
demographic collection and the second phase consists of a controlled selection
of the sample.
Oct. 12, 2005
Jaime Brugueras :
3 p.m. in SEO 512
Abstract
Markov decision processes (MDP) are stochastic processes that describe
the evolution of dynamic systems controlled by sequences of decisions or
actions. Different paths of the system lead to associated economic consequences;
the ultimate aim is to take those actions that optimize a certain criterion. I
will review the mathematical model of such processes, give real-life examples,
and describe the well-known algorithms for finding optimal policies. Stochastic
games are a natural generalization of MDP to the case of two or more
controllers. Existence of finite algorithms for finding optimal stationary
policies is in general an open problem. I consider a special class of stochastic
games, those with perfect information, which can be solved via a finite algorithm.
Oct. 19, 2005
Dr. William R. Porter :
3:30 p.m. in SEO 512
Abstract
Accelerated testing of product quality attributes is an important tool used in
the development of new products. Many products fail to meet quality
specifications after storage for prolonged times due to degradation of the
materials used to manufacture the product. The time to failure is the shelf
life of the product. Thermal stress is often used as an accelerant to promote
degradation in a controlled manner to predict shelf life. The results of stress
degradation experiments are evaluated using the nonlinear Arrhenius kinetic
model. Historical methods of fitting the Arrhenius model to thermal degradation
experimental data, particularly the Garrett approach, will be discussed.
Optimal design of thermal degradation experiments from a theoretical basis has
been described in the literature but many practical pitfalls remain. Nonlinear
model design requires knowledge of the model parameters, crude estimates of
which can be obtained using a pilot experiment. Monte Carlo evaluation of the
design of a pilot experiment will be discussed, along with opportunities for
future development of practical designs for refining parameter estimates.
Oct. 26, 2005
Wenting, Wu :
3:30 p.m. in SEO 512
Abstract
Assessing agreement plays an important role in assessing the
acceptability of a new or generic process, methodology, and formulation in
areas of laboratory performance, instrument or assay validation, method
comparisons, and individual bioequivalence. An introduction of current
existing methods for measuring agreement will be given. After that, an unified
approach for measuring agreement of k readers each with multiple readings will
be introduced. When each reader has only one reading, this mehtod degenerates
into the conventional overall agreement index. When there are only two
readers, it degenerates into the conventional concordance correlation
coefficient (CCC). When data are ordinal, it degenerates into the weighted
kappa coefficient. When data are binary, it degenerates into the kappa
coefficient. The approach uses generalized estimation equations (GEE) to
model various functions of variance components. Some simulation results will
be presented and followed by an example.
Nov. 2, 2005
Mohsen Pourahmadi :
3:30 p.m. in SEO 512
Abstract
We survey the progress made in modelling covariance matrices from the
perspective of generalized linear models (GLM) and show how one can move beyond
the use of the identity and logarithmic link functions, and prespecified
structures. Observing that most time-domain models (ARMA, state-space,....) in
time series analysis are means to diagonalize a Toeplitz covariance matrix via a
unit lower triangular matrix (Cholesky decomposition), we discuss the
distinguished role of the Cholesky decomposition in providing a systematic and
data-based procedure for formulating and fitting parsimonious models for general
covariance matrices guaranteeing the positive-definiteness of the estimates.
Pulling together some techniques from regression and time series analyses
provide the necessary tools for the procedure which reduces the unintuitive task
of modelling covariance matrices to that of a sequence of regression models. The
procedure is illustrated using a real longitudinal dataset.Once a bona fide
GLM framework for modelling covariances is found, its bayesian, nonparametric,
generalized additive and other extensions can be developed in direct analogy
with the respective extensions of the traditional GLM.
Nov. 9, 2005
Rita Saha Ray :
3:30 p.m. in SEO 512
Abstract
A critical set consists of the minimum information needed to recreate a
combinatorial structure uniquely. To date, very few results on critical sets for
a set of Mutually Orthogonal Latin Squares [MOLS] are known. In the present
talk, we consider k Mutually Orthogonal Cyclic Latin Squares of order n, n odd,
and obtain bounds on the possible sizes of the minimal critical sets. For n =
7, we consider a complete set of MOLS and exhibit
a minimal critical set, improving upon the bound reported in Keedwell (1997).
The problem is also addressed for a pair of MOLS of odd order n,
n ≥ 9. Critical sets achieving the proposed bound are obtained for n = 9 and 15.
(This is a joint work with Avishek Adhikari and Jennifer Seberry)
Nov. 16, 2005
Dr. David LeBlond :
3:30 p.m. in SEO 512
Abstract
This talk includes about 3 topics. 1. Predicting manufacturing failure rates
(one way random modeling); 2. Estimation of shelf life from accelerated
stability studies, a followup to Bill Porter's talk (non linear modeling); 3.
Comparison of 2 analytical methods, both subject to error (Errors in variables/
latent variables modeling). Also, a little of WinBUGS demo will be given.
Nov. 23, 2005
Gang Shi :
3:30 p.m. in SEO 512
Abstract
We present a statistical framework for the fixed-frequency computational
time-reversal imaging problem assuming point scatterers in a known background
medium. Our statistical measurement models are based on the physical models of
the multistatic response matrix, the distorted wave Born approximation and
Foldy-Lax multiple scattering models. We develop maximum likelihood (ML)
estimators of the locations and reflection parameters of the scatterers. Using a
simplified single-scatterer model, we also propose a likelihood time-reversal
imaging technique which is suboptimal but computationally efficient and can be
used to initialize the ML estimation. We generalize the fixed-frequency
likelihood imaging to multiple frequencies, and demonstrate its effectiveness in
resolving the grating lobes of a sparse array. This enables to achieve high
resolution by deploying a large-aperture array consisting of a small number of
antennas
while avoiding spatial ambiguity. Numerical and experimental examples are used
to illustrate the applicability of our results.
Nov. 30, 2005
Weiya Zhang :
3:30 p.m. in SEO 512
Abstract
Historically, this problem has been encountered in statistical quality
control situations and has since been extensively studied in statistical
literature on theory & applications. Notable contributors in this fascinating
topic are Wolfowitz (AMS, 1957), Searls (1963, JASA1964), Hendricks (JASA,
1964), Khan (JASA, 1968), Azen & Reed (Technoterics, 1973), Gleser & Heely
(JASA, 1976), Sen (BJ, 1979), Soofi & Gokhale (CSDA, 1991), Guo & Pal (CSAB,
2003), Chaturvedi & Tomer (Statistics, 2003) and Singh & Mathur (JSPI,
2005). Recently we came across a Review Article by Anis (2005). I will present
the review article and supplement the study with some of our own research
findings.
Jan. 18, 2006
Professor Kin-Yee Chan :
3:30 p.m. in SEO 512
Abstract
Logistic regression is a powerful technique for fitting models to data with a
binary response variable, but the models are difficult to interpret if
collinearity, nonlinearity, or interactions are present. Besides, it is hard to
judge model adequacy since there are few diagnostics for choosing variable
transformations and no true goodness-of-fit test. To overcome these problems,
we propose to fit a piecewise (simple,multiple or stepwise) linear logistic
regression model by recursively partitioning the data and fitting a different
logistic regression in each partition. This allows nonlinear
features of the data to be modeled without requiring variable transformations.
Trend-adjusted chi-square tests are used to control bias in variable selection
at the intermediate nodes. This protects the integrity of inferences drawn from
the tree structure. The binary tree that results from the partitioning process
is pruned to minimize a cross-validation estimate of the predicted deviance.
This obviates the need for a formal goodness-of-fit test. Our algorithm, called
"LOTUS", is compared with standard stepwise logistic regression and two
well-known classification tree algorithms (QUEST and C4.5) on 13 real
datasets, with several containing tens to hundreds of thousands of observations.
Results will be presented at this talk.
Feb. 6, 2006
Professor Probal Chaudhuri :
4 p.m. in SEO 636
Abstract
Statistical learning problems arise in the study of molecular evolution as well
as in the development of gene predictors. Problems can be unsupervised,
supervised or partially supervised in nature depending on the situation. This
talk will discuss the use of oligonucleotide distributions in DNA sequences in
solving such problems. Some related probabilistic models for DNA sequences will
also be discussed. In particular, a stochastic replication model for biological
sequences will be introduced that generalizes standard hidden Markov models.
March 1, 2006
Dr. Jeen Liu :
3:30 p.m. in SEO 512
Abstract
This presentation will be an overview of the role of statistics and
statisticians in the drug development process. I will present an overview of
the pharmaceutical industry, drug development process, and the clinical research
organizations. That will be followed by the challenges for statisticians in the
industry. The current statistical issues of interest and some examples will
also be provided.
March 8, 2006
Dr. Mahtab Munshi :
3:30 p.m. in SEO 512
Abstract
We examine the impact of missing data in two settings, the development of
prognostic models and the addition of new risk factors to existing risk
functions. Most statistical software presently available performs complete case
analysis, wherein only participants with known values for all of the
characteristics being analyzed are included in model development. Missing data
also impacts the summarization of evidence amongst multiple studies using
meta-analytic techniques. As we progress in medical research, new covariates
become available for studying various outcomes. While we want to investigate the
influence of new factors on the outcome, we also do not want to discard the
historical datasets that do not have information about these markers. We
investigate different methods to estimate parameters for a model when some of
the covariates are missing. These methods include likelihood-based inference for
the study-level coefficients and likelihood based inference for the logistic
model on the person-level data. We compare the results from our methods to the
corresponding results from complete case analysis. We focus our empirical
investigation on a historical example, the addition of high-density lipoproteins
to existing equations for predicting death due to coronary heart disease. We
verify our methods through simulation studies on this example.
March 15, 2006
Weiya Zhang, Ph.D. Candidate :
2:30 p.m. in SEO 512
Abstract
Inference on a normal mean with known CV is intricate since the distribution
does not admit of a complete sufficient statistic. Consequently, no umvue exists
for the mean. As a result, there have been many attempts to suggest biased
estimators which are functions of the sample mean and the sample sd. There is an
extensive literature on this fascinating topic. However, the mean parameter is
assumed to be positive-valued. If we allow the parameter space to include
negative values of the mean as well, the problem of
unbiased estimation of the mean in terms of the sample standard deviation
becomes intriguing. Starting with the framework of n(>1) i.i.d. observations
from a normal
population with known CV, we offer (i) an analytical expression for exact
unbiased estimator of the mean in terms of the sample sd and the sign function
of the sample mean;
(ii) an analytical expression for the best linear combination of the estimator
in (i) and the sample mean as an unbiased estimator for the mean, along with
exact expression for the variance of this linear combination; (iii) a study of
asymptotic normality of the linear combination in (ii) as well as its behavior
in small samples; (iv) confidence interval for the mean based on a variation of
the best linear combination in (ii); (v) comparison of traditional confidence
interval for the mean and the one suggested in (iv); (vi) improved fixed width
confidence interval for the mean.
March 29, 2006
Dr. Weining Z Robieson :
3:30 p.m. in SEO 512
Abstract
Clinical trial design issues from statistical point of view will be presented.
Statisticians' role in a pharmaceutical company will be illustrated by
description of statisticians' job responsibilities. The following issues will be
discussed: 1)Handling of missing data; 2)Adjustment for baseline covariates;
3)Interim analysis; 4)Multiple comparison; 5)Adaptive design.
Professor Min Yang :
10 a.m. in SEO 512
Abstract
Crossover designs, where experimental subjects are used in two or more (p)
periods for the purpose of evaluating and studying two or more (t) treatments,
originated from agricultural studies and have proven widely effective in a
variety of fields, especially in phase I and phase II pharmaceutical clinical
trials. The rigorous study of these designs and their optimality and efficiency
has a history of more than 3 decades. In this talk, we will review and study the
optimality, efficiency, and robustness of crossover designs under the following
two different situations (i) all treatment comparisons are equally important and
(ii) for comparing several test treatments to a control treatment. Two
algorithms, both guided by these efficiencies and results from optimal design
theory, are proposed for obtaining efficient designs under the various models.
April 5, 2006
Dr. Billy Franks :
3:30 p.m. in SEO 512
Abstract
Many characteristics for predicting death due to coronary heart disease are
measured on a continuous scale. These characteristics, however, are often
categorized for clinical use. We suggest a systematic approach to determine the
best categorizations of systolic blood pressure and cholesterol level for use in
identifying individuals who are at high risk for death due to coronary heart
disease. We also compare these data derived categories to those in common usage.
A version of Classification And Regression Trees (CART) that can be applied to
censored survival data will be used to identify categories in multiple data
sets. The collection of categories will then be used to identify major
cut-points, which are common in all of the data sets by using kernel density
estimation.
April 12, 2006
Professor Dulal K. Bhaumik :
3:30 p.m. in SEO 512
Abstract
Modern methods for imaging the human brain, such as functional magnetic
resonance imaging (fMRI)present a range of challenging statistical problems. In
this talk, we will look at a number of possible models for the analysis of fMRI
data from multiple subjects, and develop tests that lead to calculations of
power and sample size for between group comparisons. Sample size calculations
are particularly critical for neuroscientists who use these new techniques,
since each subject is expensive to image.
April 14, 2006
Yuping Dong :
3:30 p.m. in SEO 512
Abstract
The exponentially weighted moving average methods (EWMA) are one
of the statistical surveillance methods commonly studied in statistical
process control literature. The EWMA methods are mainly used to monitor
the mean of the distribution of a continuous quality measure. In this
talk, we present a way to extend the EWMA procedure to the case of a
positive shift in the incidence rate per exposure unit of a Poisson
process. Three types of EWMA methods, EWMAe, EWMAa1 and EWMAa2, are
constructed, all with an alarm statistic, which is an exponentially
weighted moving average of the observations per exposure unit. Analytical
bounds for different measures of evaluation, suitable in different types
of applications, are provided such as the expected delay, the average run
length to an alarm and the probability of successful detection, to give a
broad picture of the features of the methods. Results from a simulation
study are presented both for a fixed average run length to the first false
alarm and a fixed probability of a false alarm.
April 26, 2006
Gib Bassett :
3:30 p.m. in SEO 512
Abstract
After a brief introduction to quantile regression and modeling the talk
will consider paired comparisons. The context is rating and ranking sports
teams based on game outcomes. The quantile regression approach includes the
standard model as a special case, while allowing for a richer set of
possible relationships between teams and outcomes. Compared to models that
focus on one part of a distribution, the quantile approach expresses
relationships that depend on the different parts of a distribution. We can
have "A better than B" based on the expected outcome, while at the same
time B is more likely to win the game. The ratings are defined as handicaps
that make handicap-adjusted outcomes of games evenly matched (where
"equally matched" depends on which property of the outcome distribution is
to be equalized). We consider connections to point spreads and odds
wagering as well as the roundness of a round robin.
Sept. 20, 2006
Peter McCullagh :
3:30 p.m. in SEO 636
Abstract
Trees arise naturally in the study of the evolution of a
population. Branching processes are the natural tool in the forward
direction, and coalescent processes are natural for the study of ancestral
relationships or lineages in reverse time.
A tree has a natural graphical representation, but for some purposes a
matrix representation is also useful.
In statistical work, a similarity matrix is a covariance matrix generated
by additive common factors with independent components. The set of
similarity matrices also coincides with the set of fragmentation trees.
Some issues arising in the use of structured covariance matrices of this
sort will be discussed.
Sept. 27, 2006
S. Hedayat :
3:30 p.m. in SEO 512
Oct. 4, 2006
Tonglin Zhang :
3:30 p.m. in SEO 512
Abstract
Moran's I is the most widely used and the most frequently cited test statistic
in spatial statistical literature. This research bridges the permutation test of
Moran's I to the residuals of a loglinear model under the asymptotic normality
assumption. It provides the versions of Moran's I based on Pearson residuals ( )
and deviance residuals ( ) so that they can be used to test for spatial
clustering while at the same time account for potential covariates and
heterogeneous population sizes. Our simulations showed that both and are
effective to account for heterogeneous population sizes. The tests based on
and are applied to a set of loglieanr models for early stage and late-stage
breast cancer with socioeconomic and access-to-care data in Kentucky. The
results showed that socioeconomic and access-to-care variables can sufficiently
explain spatial clustering of early stage breast carcinomas, but these factors
cannot explain that for the late-stage. For this reason, we used local spatial
association terms and located four late-stage breast cancer clusters that could
not be explained. The results also confirmed our expectation that a high
screening level would be associated with a high incidence rate of early stage
disease, which in turn would reduce late-stage incidence rates.
Oct. 18, 2006
Professor Marc Hallin :
3 p.m. in SEO 636
Abstract
The modern history of ranks in statistics started in 1945 with Frank
Wilcoxon's far-reaching four page paper on rank tests for location. Emphasis in
1945 was on distribution-freeness and ease of applications. Since then, under
the impulse of such names as Chernoff, Savage, Hodges, Lehmann, Hajek,and Le
Cam, rank-based methods have followed the development of contemporary
statistics, and turned into a complete body of modern, flexible and powerful
techniques. In this talk, we show how this evolution, from distribution-
freeness to group invariance and tangent space projections, eventually may
reconcile the enemy brothers of statistics---efficiency and robustness.
Oct. 25, 2006
Dr. Grace L. Yang :
3:30 p.m. in SEO 512
Abstract
Occurrence of dead time in recording instruments poses challenging problems in
data acquisition, construction of stochastic models and statistical analysis.
Well-known examples include the construction of probability models for a
paralyzable counter (electron multiplier) and a nonparalyzable counter (e.g.,
Geiger counter). In this presentation, statistical analysis of recordings from
Phase Doppler Interferometry (PDI)is considered. PDI is a non-intrusive
technique used to obtain information about spray characteristics in many
areas of science, such as liquid fuel spray in combustion, spray coatings, fire
suppression and pesticide dispensing. PDI can record the velocity of individual
droplets in a spray. However, it will miss some of the droplets because of a
recurring presence of dead time. The incompleteness of PDI recordings results in
a multimodal interarrival time distribution of droplets. Modeling a spray
process as a homogeneous Poisson process, we estimate the spray diffusion rate
(Poisson intensity) with correction for dead time under various conditions. The
asympotic distribution of the estimates is derived from a strict stationary
process. Simulation produced a good agreement between our estimators (in the
presence of dead time) and the MLE obtained without dead time. Experimental data
from NIST are used for illustration.
Nov. 15, 2006
Professor Sanjib Basu :
3:30 p.m. in SEO 512
Abstract
The rates of cancers, including age-adjusted mortality and incidence rates,
depict a general increase over the last 30 years. These led some to
question the success of the war on cancer. The rates of many other
competing diseases, on the other hand, have declined. It has been
hypothesized that this decline is somewhat responsible for the rise in
cancer rates. We consider competing risks analysis of cancer survival data
that considers the simultaneous risks of cancer as well as other causes. The
cure rate survival models for cancer postulates a fraction of the patients to be
cured from cancer. We propose a model that incorporates competing risks and, at
the same time, allows a fraction of patients to be cured. We describe Bayesian
analysis of this model, discuss both conceptual and methodological issues
related to model building and model selection, and consider application in
survival data for breast and prostate cancer patients in the SEER registries of
the National Cancer Institute (NCI).
Nov. 29, 2006
Li Wei :
3:30 p.m. in SEO 512
Abstract
Stochastic curtailment, one of the major statistical tools adopted in interim
analysis, has attracted more attentions than its competitors such as group
sequential procedures for it integrates current data and potential future
outcomes in addition to its simplicity in design and implementation. Under this
approach, the conditional power, which is the probability of rejecting the null
hypothesis at the planned end of the study given the accumulating data, is
calculated and the stopping decision is made according to the comparison of this
power with a pre-specified threshold. Many procedures with this perspective have
been developed for interim analysis. However, possibly for the purpose of
statistical convenience, only trials with one or two arms are investigated.
Here, we derived an analytic formula for the conditional power under the frame
of linear models so that it can be applied to most actual clinical trials in
which multiple treatment effects, block effects and covariate effects are all
allowed to be considered. The properties of this conditional power is
investigated and further our research shows that, unlike the standard power of a
regular test for a treatment contrast which depends on unknown parameters only
through the contrast itself, the conditional power here fails to have this
characteristic in general. A necessary and sufficient condition for the
conditional power to depend soly on the interested contrast is provided and some
instances are illustrated. Similar arguments can be made about the sufficient
statistics for the conditional power. Finally, the results obtained here is
applied to an interim analysis performed in a multi-center, randomized,
double-blinded, placebo-controlled, parallel group phase II study where centers
act as blocks and baseline scores are treated as covariates, resulting in an
early termination of the trial and hence a substantial saving in cost.
Jan. 31, 2007
Prof. Klaus Miescke :
4 p.m. in SEO 712
Abstract
To estimate the total overpayment on a large number of paid claims, government
agencies utilize extrapolation methods that are based on the audit results from
a random sample of these claims. Although there is a variety of approaches that
are reasonable and statistically valid, some play out better than others in the
legal process. For example, the estimator of the total loss must be unbiased by
fairness reasons. Using a biased estimator with a smaller MSE would be
unacceptable. In this talk we describe the entire audit process, including
possible objections from the other side and suggestions on how to respond to
them. The theoretical part of the talk is based on W.G. Cochran (1977), Sampling
Techniques, 3rd ed., Wiley, NY, and the rest on practical experience. It should
be pointed out that it is the loss, not the fraud itself, that can be detected
and estimated
statistically.
Feb. 14, 2007
Jie Yang :
3:30 p.m. in SEO 712
Abstract
We propose a new class of dimension reduction methods
using the first two inverse moments, called Sliced Inverse Moment
Regression (SIMR). We develop corresponding weighted chi-squared
tests for the dimension of the regression. Basically, SIMR are
linear combinations of Sliced Inverse Regression (SIR) and a new
method using candidate matrix M_{zz'|y}, which is designed to
recover the entire inverse second moment subspace. Theoretically,
SIMR, as well as Sliced Average Variance Estimate (SAVE), are more
capable of recovering the complete central dimension reduction subspace
than SIR and Principle Hessian Directions (pHd). Therefore it can
substitute for SIR, pHd, SAVE or any linear combination of them
at a theoretical level. Simulation study shows that SIMR using the
weighted chi-squared test may have consistently greater power than
SIR, pHd, and SAVE.
March 5, 2007
Prof. Emad Aly :
3:30 p.m. in SEO 712
Abstract
We consider the problem of testing the null hypothesis of no change against the
alternative of multiple change points in a series of independent observations.
We consider the three cases of testing against the general multiple change point
alternative, the ordered multiple change point alternative and the epidemic two
change point alternative. We report the asymptotic null distribution of the
considered tests. We also give approximations for their limiting critical values.
March 7, 2007
Prof. Lijian Yang :
3:30 p.m. in SEO 712
Abstract
For the past two decades, single-index model, a special case of projection
pursuit regression, has proven to be an efficient way of coping with the high
dimensional problem in nonparametric regression. Applications of single-index
model lie in a variety of fields, such as discrete choice analysis in
econometrics and dose-response models in biometrics, where high-dimensional
regression models are often employed. We investigate the single-index prediction
based on weakly dependent sample. The single-index is identified by the best
approximation to the multivariate prediction function of the response variable,
regardless of whether the prediction function is a genuine single-index
function. A polynomial spline estimator is proposed for the single-index
coefficients, and is shown to be strongly consistent and asymptotically normal.
An iterative program based on Newton-Raphson algorithm is developed. The
algorithm is sufficiently fast for the user to analyze large data of high
dimension within seconds. Simulation experiments have provided strong evidence
that corroborates with the asymptotic theory. Finally, we illustrate our
estimation procedure by a gas furnace example.
March 9, 2007
Thomas E. Bradstreet, Ph.D. :
2 p.m. in SEO 712
Abstract
This example driven presentation outlines some of the basic types of studies
traditionally conducted in the pharmaceutical industry. Both animal and human
research activities are presented. Special attention is given to statistical
analysis issues. Some brief study descriptions and their corresponding data
sets can be found at http://www.math.iup.edu/~tshort/Bradstreet. Students in
the Friday afternoon Statistics in Medicine class can choose individual homework
assignments based upon these data sets. Further opportunities include writing
data set based teaching papers, and potential research topics. In addition, a
high level introduction to newer areas of statistical initiatives will be provided.
March 21, 2007
Prof. Gabor J. Szekely :
3:30 p.m. in SEO 712
Abstract
We introduce a simple new measure of dependence between random vectors. Distance
covariance (dCov) and distance correlation (dCor) are analogous to
product-moment covariance and correlation, but unlike the classical definition
of correlation, dCor = 0 characterizes independence for the general case. The
empirical dCov and dCor are based on certain Euclidean distances between sample
elements rather than sample moments, yet have a compact representation analogous
to the classical covariance and correlation. Definitions can be extended to
metric-space-valued observations where the random vectors could even be in
different metric spaces. Asymptotic properties and applications in testing
independence will also be discussed. A new universally consistent test of
multivariate independence is developed. Implementation of the test and Monte
Carlo results are presented.
April 18, 2007
Professor Per Mykland :
3:30 p.m. in SEO 712
Abstract
In the econometric literature of high frequency data, it is often assumed that
one can carry out inference conditionally on the underlying volatility
processes. In other words, conditionally Gaussian systems are
considered. This is often referred to as the assumption of ``no leverage
effect". This is often a reasonable thing to do, as general estimators and
results can often be conjectured from considering the conditionally
Gaussian case. The purpose of this paper is to try to give some more structure
to the things one can do with the Gaussian assumption. We shall argue in the
following that there is a whole treasure chest of tools that
can be brought to bear on high frequency data problems in this case. We shall in
particular consider approximations involving locally constant volatility
processes, and develop a general theory for this approximation. As applications
of the theory, we propose an improved estimator of quarticity, an ANOVA for
processes with multiple regressors, and an estimator for error bars on the
Hayashi-Yoshida estimator of quadratic covariation.
April 20, 2007
Professor Ajit C. Tamhane :
2 p.m. in SEO 512
Abstract
This talk is in two parts. Part I will give a brief introduction to multiple
comparison procedures to provide the necessary background for Part II which is
the main topic. For brevity and simplicity, we shall restrict to procedures
based on p-values only.
Parallel and serial gatekeeping procedures have been recently proposed (Westfall
and Krishen 2001, and Dmitrienko, Offen and Westfall 2003) for testing
hierarchically ordered families of hypotheses. We generalize these procedures to
what we call tree-structured gatekeeping procedures. This generalization is
necessary to deal with problems involving hierarchically ordered multiple
objectives subject to logical restrictions, e.g., in the analysis of multiple
endpoints in dose-control studies and in superiority-equivalence testing. The
proposed approach is based on the closure principle of Marcus, Peritz and
Gabriel (1976) and uses weighted Bonferroni tests for intersection hypotheses.
In special cases of parallel or serial tree structures the closed testing
procedure can be shown to be equivalent to stepwise procedures, which are easy
to implement. Two illustrative clinical trial examples are given.
Note: This work is joint with Alex Dmitrienko, Brian Wiens and Xin Wang, and is
based on a paper that has recently appeared in Statistics in Medicine.
April 25, 2007
Professor Vijay Nair :
3:30 p.m. in SEO 712
Abstract
The term network tomography characterizes two classes of large-scale inverse
problems that arise in the modeling and analysis of computer and communications
networks. One class of problems deals with passive tomography where network
traffic data are collected at the nodes, and the goal is to reconstruct
origin-destination traffic patterns. The second one is active network tomography
where the goal is to recover link-level quality of service parameters, such as
packet loss rates and delay distributions, from end-to-end path-level
measurements. Internet service providers use this to characterize network
performance and to monitor service quality. This talk will provide an overview
of the network application, the statistical inverse problems that arise, and
some recent research in trying to address them with an emphasis on active
tomography. This is joint work with George Michailidis, Earl Lawrence, Bowei Xi,
and Xiaodong Yang.
April 27, 2007
Professor Teresa Azinheira Olivira :
2 p.m. in SEO 512
Abstract
The Fisher related information of a balanced block design will remain invariant
whether or not the design has repeated blocks. This fact can be used
theoretically to build a large number of non-isomorphic designs for the same set
of design parameters. Further, as it has been shown by UIC researchers and many
others later designs with repeated blocks could be used for many different
purposes both in experimentations and surveys from finite populations. In this
talk the subject will be briefly reviewed and new results on the existence and
construction of such designs will be presented. Several unsolved problems for
further research will be presented.
May 2, 2007
Hongmei Liu and Nordia Thomas :
3:15 p.m. in SEO 712
Abstract
The objective of our project was to evaluate the prevalence of
math anxiety for two periods (A and B) of Algebra I students, and to
identify if alternative pedagogical strategies help to alleviate math
anxiety. From an initial Math Anxiety Scale (MAS) survey we concluded
that students did exhibit math anxiety, with most displaying medium to
medium-high levels of math anxiety. We conducted a logistic regression of
the difference of scores between the Math Anxiety Rating Scale (MARS)
survey given before and after a class, and the students' initial math
anxiety rating. Using the results of the logistic regression we concluded
that the hands-on teaching approach used in Period A helped to alleviate
the students' math anxiety more than the traditional lecture approach
employed in Period B.
Sept. 5, 2007
Prof. Michael Stein :
3:30 p.m. in SEO 636
Abstract
For Gaussian spatial processes observed at a large number of irregularly sited
locations, exact calculation of the likelihood is generally not possible due to
both memory and computational constraints. If we can write the covariance matrix
of the observations as a sparse matrix plus a matrix of moderate rank, then both
the number of computations and memory requirements can be greatly reduced. The
idea is that the sparse term will capture the local behavior of the process and
the low rank term the large-scale behavior. This approach is applied to compute
likelihood-based estimates of the spatial covariance structure for total column
ozone measurements on a global scale. The approach can be judged a success
computationally in that likelihoods can be calculated exactly for datasets far
too large to carry out the computations for a more general model. However,
various diagnostics show problems with the model, so that further work is needed.
Sept. 12, 2007
Shi Zhao, PhD Candidate :
3:30 p.m. in SEO 712
Abstract
Latin square constructed crossover designs balanced for first order
carryover effects, commonly referred to as Williams designs, are commonly
used in many PK, PD, and other clinical studies. These designs are used to
investigate treatment (t) effects in the presence of two nuisance factors,
subjects (s) and periods (p), while evaluating and accommodating first
order carryover effects with equal precision among treatment comparisons. In
some studies, an additional design factor and higher order carryover
effects are of interest. For example, in capsaicin cough challenge studies, the
additional design factor cough counter (c) is of interest, as is the
possibility of higher order carryover effects given the nature of the
capsaicin induced cough endpoint, and the human interaction between the
cough counters and the study subjects.
Graeco-Latin square crossover designs balanced for up to t-1 and c-1 order
residual effects need to be constructed. We illustrate for the case where t = p
= c = 4. Specifically, two sets of three 4x4 mutually orthogonal Latin square
(MOLS) crossover designs with 12 sequences are constructed. Then
permutations of each set of three 4x4 MOLS are enumerated. Particular
permutations from the first set of three 4x4 MOLS are selected and
superimposed upon selected permutations from the second set of three 4x4
MOLS, to construct the desired Graeco-Latin square crossover design with 24
sequences of treatment and evaluator combinations. Using field theory, this
result is generalized to any case where t is prime number or a power of a
prime number. Open questions related to these results such as: how to
partially balance the treatment and evaluator combinations across periods;
how to construct designs for numbers of treatments which are not primes or
powers of primes; will also be discussed. Appropriate intellectual building
blocks will be presented throughout the talk.
Oct. 3, 2007
Jing Wang :
3:30 p.m. in SEO 712
Abstract
Asymptotically exact and conservative confidence bands are obtained for
nonparametric regression function, based on piecewise constant and piecewise
linear spline estimation, respectively. Compared to the pointwise nonparametric
confidence interval of Huang(2003), the confidence bands are inflated only by a
factor of log(n)^{1/2}, similar to the Nadaraya-Watson confidence bands of
Hrdle(1989), and the local polynomial bands of Xia(1998) and Claeskens andVan
Keilegom(2003). Simulation experiments have provided strong evidence that
corroborates with the asymptotic theory. Testing against the linear spline
confidence band, the commonly used trigonometric trend is rejected with highly
significant evidence for the Leaf Area Index of Aquatic Agriculture land, based
on the remote sensing data collected from East Africa.
Oct. 10, 2007
Leping YIn and Wei Zheng :
3:30 p.m. in SEO 712
Abstract
The client's objective is to find important factors causing the patients' voice
handicap and to investigate the influence of those factors. There are several
potential factors of interest: Diagnosis (kind of disease), Gender, Ethnicity,
voice therapy (whether the patient complies with doctors ~ R order or not), Age,
and Singer or not. All the explanatory factors are categorical variables. We run
PROC GLM in SAS and fit an ANOVA model with interactions between the factors.
The following factors are found significant in the data analysis: Diagnosis,
Singer, Therapy, Gender, and Gender*Diagnosis.
Oct. 17, 2007
Dr. Mohammad F. Huque :
3:30 p.m. in SEO 712
Abstract
In evaluating that a test treatment is safe and effective in treating a disease,
it is often necessary in clinical trials to answer more than one clinically
relevant question or to characterize a treatment effect in two or more
endpoints. This requires framing clinically relevant multiple hypotheses
involving multiple primary and secondary endpoints and treatment comparisons.
These multiple hypotheses can be statistically tested according to a strategy
that depends on the objectives of the trial and clinical considerations of the
disease and the treatment under study. However, such a statistical testing
strategy concerning multiple hypotheses involving multiple endpoints is fraught
with multiplicity issues, which if ignored can increase the chance of spurious
positive findings resulting in false inferences that an effect is shown when
there is really no such an effect. In interpreting results from a situation like
this, one is more likely to make a false conclusion about the benefit of the
test treatment because there are multiple opportunities to choose favorable
results from multiple analyses. Therefore, it is necessary that a study protocol
of the trial include a clear plan for addressing multiplicity issues and a
statistical testing strategy that is appropriate for a given benefit claim of
the study treatment. This presentation will examine some basic principles and
statistical considerations that can be helpful in better planning and testing of
multiple endpoint hypotheses in clinical trials.
Oct. 24, 2007
Prof. Douglas G. Simpson :
3:30 p.m. in SEO 712
Abstract
A class of heteroscedastic generalized linear regression models is developed in
which a subset of the regression parameters are scaled nonparametrically.
Efficient semiparametric inferences are derived for the parametric components of
the models. Bootstrap tests for scale heterogenerity are also developed. The
models provide an approach to adapt for heterogeneity in the data due factors
such as to varying exposures and varying levels of aggregation. The methodology
is illustrated with simulations, published data and data from collaborative
research on ultrasound safety.
Oct. 31, 2007
Prof. Stephen Portnoy :
3:30 p.m. in SEO 712
Abstract
In many situations where censored observations are observed, it is not
unreasonable to assume that the censoring values are known for all observations
(even the uncensored ones). For example, one of the earliest approaches to
ensored regression quantiles was introduced by work of Powell in the mid 1980's.
Powell assumed that the censoring values were constant, thus positing
observations of the form Y = min(T, c) (where Y is observed and T is the
possibly unobserved survival time that is assumed to obey some linear model).
More generally, we may be willing to assume that we observe a sample of
censoring times {ci} and a sample of censored responses Yi = min(Ti, ci) , a
model that could apply to a single sample. In this case, one could use the
empirical distributions of the {Yi} and {ci} and take the ratio of empirical
survival functions to estimate the survival function of T. This is
asymptotically equivalent to applying the Powell method on a single sample.
Despite some optimality claims of Newey and Powell, it turns out that the Kaplan
Meier estimate is better (asymptotically, and by simulations in finite samples)
even though it does not use the full sample of {ci} values. More generally, even
in multiple regression settings, the censored regression quantile estimators
(Portnoy, JASA, 2003) are better in simulations than Powell's estimator (even
for the constant censoring situation for which Powell's estimator was
developed). Remarkably, in the one sample case, replacing the empirical function
of {ci} by the true survival function (assuming it is known) yields an even less
efficient estimator. Thus, it appears that discarding what appears to be
pertinent information improves the estimators. The talk will try to quantify and
explain this conundrum.
Nov. 7, 2007
Dr. Kooros Mahjoob :
3:30 p.m. in SEO 712
Abstract
In some randomized clinical trials, missing values arise due to patients'
discontinuation before the end of the trial. As a result, there will be no
value/measurement for those patients who dropped out for assessing efficacy at
the end of the trial. Such a phenomenon is typical in neurological and
psychiatric clinical trials; in fact, in some cases, there are over 40%
dropouts. Clearly, analyzing trials data sets that have missing values and then
drawing conclusions from the analysis results is a challenging task for FDA
statisticians. Dealing with missing values in clinical trials has a long
history, which goes back for over two decades. Lots of work, published papers
and technical notes have suggested methods to deal with the issue and have
talked about the utility of one method over the others. Nevertheless, the
reality is that there is no clear-cut solution to the problem. Often, in some
trials, missing values is a real problem from the regulatory perspective as to
how to make a decision on the drug approval.
This presentation will focus on framing the problem, discussing common
statistical methods used and the FDA's views on the methods. I will also discuss
the result of simulations, bootstrapping using data of actual trials, conducted
by FDA colleagues, in comparing the performances of different methods, and the
outlines of some new methods.
Nov. 14, 2007
Michael Levine :
3:30 p.m. in SEO 712
Abstract
We investigate several possible strategies for consistently estimating the
so-called Hurst parameter H responsible for the long-memory property in a
special class of nonlinear ARCH-type models popularly known as LARCH, as well as
in the continuous-time Gaussian stochastic process named fractional Brownian
motion (fBm). Several estimation methods are discussed, including a conditional
MLE method and a local Whittle-type estimation procedure. The conditional MLE is
proved to be consistent and a Portmanteau-type test for model validation is
established. By constructing the LARCH and fBm processes on a common probability
space, and showing the convergence of various partial sums of the former to the
latter in mean squared, we can propose a specially designed conditional maximum
likelihood method for estimating the fBm's Hurst parameter. In keeping with the
popular financial interpretation of ARCH-type models, all estimators are
based only on observation of the "returns" of the model and not on the
"volatilities".
Nov. 16, 2007
Yuan Xu :
2 p.m. in SEO 712
Abstract
Patients with heart disease usually stay in the hospital for certain days before they are cured and then leave the hospital. For each patient, the amount of money that he/she spends in the hospital usually is different. Some people pay more, some the other people pay less. Now we have total 143211 records of such patients. For each patient's record, we have his/her total expenditure during stay in hospital and a lot of other information such as his/her sex, age, race, length of stay in hospital, different disease history for example once having shock or not, different hospital conditions, different financial conditions for example insurance and so on. It is observed that in average Hispanic Women or Hispanic Men tend to spend more money during their stay in hospital than other race group of people. Our objective is to give a good explanation on what is the cause for the above phenomenon.
After discussion with medical doctors who provided valuable suggestions, we started with 25 suspected variables. After running the linear regressions for the total and each of the 8 sub-groups of patients, we excluded many of these variables and concentrated our investigation on 5 left variables. We applied different kinds of statistical tests on these 5 variables and found that most likely only two of them may contribute to the higher cost of Hispanic patients. By medical doctor's opinion, we excluded one of these two variables and kept the last one as our investigation result. We claim that the special behavior of Hispanic patients regarding to this variable actually causes them to pay more in the hospital.
Nov. 28, 2007
Weibiao Wu :
3:30 p.m. in SEO 712
Abstract
I will talk about statistical inference of trends in mean non-stationary models,
and mean regression and conditional variance (or volatility) functions in
nonlinear stochastic regression models. Simultaneous confidence bands are
constructed and the coverage probabilities are shown to be asymptotically
correct. The Simultaneous confidence bands are useful for model specification
problems in nonlinear time series. The results are applied to environmental and
financial data-sets.
Jan. 23, 2008
Prof. Qizhi Chen :
3:30 p.m. in SEO 712
Abstract
This paper considers the effect of imperfect vaccination in a susceptible infected removal (SIR) epidemic model. The minimum proportion of the population that needs to be vaccinated to prevent a major epidemic depends on the vaccine efficacy and the basic reproductive rate for the SIR model, allowing for imperfect and variable vaccination. Martingale theory is used to derive estimates and associated standard errors for these parameters. Asymptotic properties of the resulting estimators are investigated. Data for a mumps outbreak are used as an illustrative example.
Jan. 29, 2008
Junhui Wang :
3 p.m. in SEO 636
Abstract
Hierarchical classification is critical to knowledge and context management as well as knowledge exploration, as in gene function classification and discovery and document categorization. In hierarchical classification, an input is classified by a structured hierarchy. In a situation as such, the central issue is how to effectively utilize inter-class relationship to improve the generalization performance of flat classification ignoring such dependency. In this talk, a novel large margin method based on constraints characterizing multi-path hierarchy is presented within the framework of regularization. In particular, I will discuss three aspects: (1) the idea and methodology development; (2) computational tools; (3) a statistical learning theory. Numerical examples will be provided to demonstrate the advantage of our proposed methodology against other existing competitors. An application to gene function prediction and discovery will be discussed.
Feb. 28, 2008
Yufei Chen and Weiyun Zheng :
1 p.m. in SEO 712
Abstract
Zenker's diverticulum is a diagnosis that primarily affects individuals
in the seventh and eighth decades. The current mainstay of treatment is
surgical. Due to the advanced age of the population, many are not prime
surgical candidates and have a higher probability of sequelae from general
anesthetics. A newer method of delivery is injections that are given in
clinic under EMG guidance. This allows patients who have comorbidities
that make them poor surgical candidates the ability to receive treatments.
To see the value of the botulium toxin injections, our study will look at
the subjective symptomatic improvement noted by patients after receiving
the injections. Patients who will receive botox injections in the clinic
setting are asked to fill out a survey about their current status before
the injections and the improvements they will notice after treatment in
eating ability, normalcy of diet, and understandability of speech.
With the distribution of patients populations to each category, we
firstly use independent Z-test to get a rough total sample number with
sufficient power. Second, we use advanced model with conditional
probability transition matrix to simulate the whole process. The two
results will be compared and the advantage of the later will be discussed.
March 5, 2008
Dr. Cong Han :
3:30 p.m. in SEO 712
Abstract
This presentation will review design issues for studies of HIV dynamics that use a nonlinear mixed-effects model. A method based on a first-order approximation, used for similar design issues in
pharmacokinetics and pharmacodynamics, will be reviewed and discussed. Limitations of this method will be discussed and other methods will be described, including methods based on the exact
calculation of the Fisher information matrix and Bayesian methods, which, although computationally intensive, provide alternatives.
March 12, 2008
Prof. Fuxia Cheng :
3:30 p.m. in SEO 712
Abstract
ARCH(p)-model has found much interest in financial econometrics. It was introduced by Engle(1982) in order to provide a framework in which so-called volatility clusters may occur, i.e., periods of high and low (conditional) variances depending on past values of the series. The model was later extended into various directions. In most of the work, the main focus has been on estimating the unknown parameters.
But it is of interest and of practical importance to know the nature of the innovation distribution. Actually, if the distribution of the innovation is unspecified, the parametric component only partly determines the distribution behavior. It is as important to investigate the distribution of the innovation as estimating the parameters. In this talk, we consider the consistency and the asymptotic distribution of the innovation density estimators in ARCH(p)-time series. We also extend the central limit theorem (CLT) and the strong law of large number (SLLN) to the average of the residuals.
March 19, 2008
Prof. Chunming Zhang :
3:30 p.m. in SEO 712
Abstract
Functional magnetic resonance imaging (fMRI) aims to locate activated regions in human brains when specific tasks are performed. The conventional tool for analyzing fMRI data applies some variant of the linear model, which is restrictive in modeling assumptions. To yield more accurate prediction of
the time-course behavior of neuronal responses, the semiparametric inference for the underlying hemodynamic response function is developed to identify
significantly activated voxels. Under mild regularity conditions, we demonstrate that a class of the proposed semiparametric test statistics,
based on the local linear estimation technique, follow chi-squared distributions under null hypotheses for a number of useful hypotheses. The
asymptotic power functions of the constructed tests are derived under the fixed and contiguous alternatives. Furthermore,
a new false discovery rate approach which incorporates spatial information of voxel-wise p-values is devised for detecting the regions of activation.
Simulation evaluations and real fMRI data application suggest that the semiparametric inference procedure provides more efficient detection of
activated brain areas than the popular imaging analysis tools.
April 2, 2008
Prof. Jie Liang :
3:30 p.m. in SEO 712
Abstract
The three dimensional structures of biomolecules such as proteins and
RNAs are the basis of their biological functions. For RNA molecule,
conformational entropy is important for stability and
folding. However, it is challenging to either measure or compute
conformational entropy associated with long loops. We develop
optimized discrete $k$-state models of RNA backbone and estimate
entropy of hairpin, bulge, internal loop, and multibranch loop of long
length using an efficient sequential Monte Carlo sampling method. The
estimated entropy indicate that the Jacobson-Stockmayer model has
large errors for bulge, internal, and multibranch loops. For protein,
we study the transition state ensemble. By generating effective
samples under various experimentally derived constraints, we
characterize the transition state ensemble (TSE) during protein
folding. As TSE is short-lived, the size and shape of conformations
in TSE have been elusive. For the protein acylphosphatase, we found
TSE has diverse conformations. In contrast to previous results, we
found overall TSE can be very different from native structure of
proteins (with RMSD>12A). To predict protein functions, we develop a
method by matching local surfaces based on estimated evolutionary
information specific to individual binding region via a Bayesian Monte
Carlo approach using a continuous-time Markov model. Our method
provides a probabilistic model which characterizes protein binding
activities that may involve multiple substrates or ligands. (Joint
work with Rong Chen, Ming Lin, Zheng Ouyang, Jeffrey Tseng, and Jian
Zhang) (please visit http://www.uic.edu/~jliang for further
information).
April 9, 2008
H.M. James Hung, PhD :
4 p.m. in SEO 636
Abstract
A topic of great interest in the recent decade is design adaptation for clinical trials using the data accumulating during the course of the trial. There are good reasons for such modification of design features. Design adaptations include sample size re-estimation, enrichment of patient population, dropping a treatment arm, etc. In this presentation I shall give an introduction of this topic and a brief overview of critiques and comments to such adaptation in the literature.
April 16, 2008
Dr. Yili Pritchett :
3:30 p.m. in SEO 712
Abstract
Path analysis, first developed in the field of genetics and actively used in sociology, refers to a modeling approach for causal relationships (Wright, 1934). In the setting of clinical research, path analysis can be a useful tool to demonstrate an independent treatment effect on a disease state, which might have causal relationships with other disease states that can be treated by the same therapy. In this presentation, the idea of using path analysis in clinical trial design will be introduced. In this approach, pre-specified causal relationships can be modeled by structural equations so that the treatment effect on the disease state of interest (direct effect) will be tested after accounting for the treatment effects on the other conditions (indirect effects). The total treatment effect can be decomposed as the sum of the direct and the indirect effects, and the statistical significance of each effect can be tested. The idea and the approach will be illustrated by concrete examples where an antidepressant was investigated for its effect on the management of different types of pain. Further statistical discussions will be given on the topics of the invariance between ordinary linear regression and standardized linear regression, and the generalization from ordinary linear model to generalized linear model.
April 23, 2008
Prof. Xin Gao :
3:30 p.m. in SEO 712
Abstract
The recent years have witnessed the increasing interest in the study of graphical models. In this talk, I will discuss two related parameter estimation and model selection problems in graphical models. The first problem is to estimate the concentration matrix of a Gaussian graphical model. We propose to estimate the concentration matrix using the penalized likelihood method with the smoothly clipped absolute deviation (SCAD) penalty. The method leads to a sparse and shrinkage estimator of the concentration matrix. Using proper choice of the regularization parameter, the proposed method automatically and consistently selects the true graphical structure and produces estimator that is as efficient as the oracle estimator. We further establish the consistency of the BIC criterion to identify the true graphical structure when used with the SCAD penalty function. The second problem is regarding the graphic model with multivariate hidden Markov structure. For such high-dimensional data with complicated dependency structure, we propose to use composite likelihood approach and especially we develop COMP-EM algorithm to perform the parameter estimation in the presence of incomplete data. The composite likelihood based information criterion was employed to select the best network structure.
April 24, 2008
Matthew J. Bourque and Xuejing Wang :
2:15 p.m. in SEO 712
Abstract
A water main bringing water into a Chicago suburb was damaged on or
about January 5, 2002. The damage was not discovered until February 11,
2004, for a total of approximately 765 days. We have daily water usage
data from 1989 through 2007, and want to determine (i) whether there is
a statistically significant increase in water flow during the period
from January 2002 until February 2004 and (ii) if there is a significant
increase, to estimate the amount of extra water flow during that time.
We use intervention analysis to fit an ARIMA model, and from the
estimates of the parameters we get the result. We will also discuss more
generally about intervention analysis.
April 30, 2008
Prof. Ejaz Ahmed :
3:30 p.m. in SEO 712
Abstract
We consider the estimation problem for the parameters of generalized linear models which may have a large collection of potential predictor variables and some of them may not have influence on the response of interest. In this situation, selecting the statistical model is always a challenging problem. In the context of two competing models, we demonstrate the relative performances of shrinkage and classical estimators based on the asymptotic analysis of quadratic risk functions. We demonstrate that the shrinkage estimator outperforms the maximum likelihood estimator uniformly. For comparison purpose, we also consider the Park and Haste type estimator (variant of lasso estimator) for generalized linear models. This comparison shows that shrinkage method performs better than the lasso type estimation method when the dimension of the restricted parameter space is large. This talk ends with real-life example showing the value of new method in practice. More, specifically, we consider South African heart disease data, which was collected on males in a heart disease high-risk region of Western Cape, South Africa.
May 1, 2008
Prof. Yazhen Wang :
3 p.m. in SEO 636
Abstract
Volatilities of asset returns are central to the theory and practice of asset pricing, portfolio allocation, and risk management. In financial
economics, there is extensive research on modeling and forecasting volatility up to the daily level based on Black-Scholes, diffusion, GARCH,
stochastic volatility models and implied volatilities from option prices. Nowadays, thanks to technological innovations, high-frequency financial data are
available for a host of different financial instruments on markets of all locations and at scales like individual bids to buy and sell, and the
full distribution of such bids. The availability of high-frequency data stimulates an upsurge interest in statistical research on better estimation of
volatility. This talk will start with a review on low-frequency financial time series and
high-frequency financial data. Then I will introduce popular realized volatility computed from high-frequency financial data and present my
work on multi-scale methods for analyzing jump and volatility variations and matrix factor models for handling large size volatility matrices.
Sept. 24, 2008
Prof. Dale Rosenthal :
4:15 p.m. in SEO 612
Abstract
The problem of classifying trades as buys or sells is examined.
I propose estimated quotes for midpoint and bid/ask tests and a modeling approach
to classification. Prevailing quotes are estimated using flexible approximations to
the distribution for delays of quotes relative to trade timestamps. Classification
is done by a generalized linear model which includes improved versions of midpoint,
tick, and bid/ask tests. The model also considers the relative strengths of these tests,
can account for market microstructure peculiarities, and allows for autocorrelations
and cross-correlations in trade direction. The correlation modeling corrects for
pseudoreplication, yielding more accurate standard errors and fixed effect estimates.
Further, the model estimates probabilities of correct classification.
The model is compared to various trade classification methods using a sample of
2,836 domestic US stocks from an unexplored, recent, and readily-available data set.
Out of sample, modeled classifications are 1-2% more accurate overall than current
methods; this improvement is consistent across dates, sectors, and locations
relative to the inside quote. For Nasdaq and NYSE stocks, 1% and 1.3% of the
improvement comes from using relative strengths of the various tests;
0.9% and 0.7% of the improvement, respectively, comes from using some form of
estimated quotes. For AMEX stocks, a 0.4% improvement is attributed to using a
lagged version of the bid/ask test. I also find indications of short- and
ultra-short-term alpha.
Oct. 1, 2008
Prof. Mathias Drton :
4:15 p.m. in SEO 612
Abstract
Many statistical models are defined in terms of polynomial constraints,
or in terms of polynomial or rational parametrizations. Such algebraic
models include, for instance, factor analysis and instrumental variable
models, latent class models, and more generally, discrete and Gaussian
graphical models with hidden variables. Statistical inference in hidden
variable models is complicated by the fact that the models' parameter
spaces are typically not smooth. This is the motivation for this talk
that considers testing a null hypothesis with singularities in algebraic
models. The focus will be on the large-sample asymptotic behavior of
likelihood ratio and Wald tests.
Oct. 8, 2008
Prof. Junhui Wang :
4:15 p.m. in SEO 612
Abstract
Large margin classifiers have proven to be effective in delivering
high predictive accuracy, particularly those focusing on the
decision boundaries and bypassing the requirement of estimating the
class probability given input for discrimination. As a result, these
classifiers may not directly yield an estimated class probability,
which is of interest itself in many real applications. In this talk,
I will present a novel method to estimate the class probability
through sequential weighted classifications, by utilizing features
of interval estimation of large margin classifiers. In particular, I
will discuss four aspects: (1) the idea and methodology development;
(2) tuning parameter selection; (3) regularization solution path;
(4) a statistical learning theory. Numerical examples will be
provided to demonstrate the advantage of our proposed methodology.
Oct. 15, 2008
Prof. Yue Yin :
4 p.m. in SEO 612
Abstract
Formative assessment was hypothesized to have a beneficial impact on
students' science achievement and conceptual change, either directly or
indirectly by enhancing motivation. We designed and embedded formatives
assessments within an inquiry science unit. Twelve middle-school science
teachers with their students were randomly assigned either to an
experimental group (N = 6), provided with embedded formative assessment,
or control group (N = 6). Teachers varied significantly as to their impact
on student motivation, achievement, and conceptual change. But the impact
of the formative assessment treatment on these outcomes was not
statistically significant. Variation in both teachers' classroom
management and the degree to which they used informal formative
assessment, regardless of group, were conjectured as possible reasons for
the absence of an overall formative assessment effect.
Oct. 22, 2008
Shi Zhao, Ph. D. :
4:15 p.m. in SEO 612
Abstract
Crossover experiments are used for comparing the responses to various different stimuli or treatments in areas ranging from psychology and human factor engineering to medical and agricultural applications. They are widely used in the pharmaceutical industry.
There is an extensive literature that assures us that a carefully designed crossover study will produce a wealth of information that will enable inference with high precision. This is based on the implicit, but critical, assumption that the experiment will yield all the planned observations. Yet in many situations, such as clinical trials, there is a substantial probability that some subjects will drop out of the study prior to the completion of their treatment sequence. Low, Lewis and Prescott (1999) observed that a dropout rate of between 5% and 10% is not uncommon and, in some areas, can be as high as 25%. They gave an example of a design in four periods based on a Williams Latin square where there is substantial loss of information if some observations are unavailable in period 4. Indeed, if all observations in the final period are not available, the design becomes disconnected, i.e., elementary contrasts are no longer all estimable. Majumdar, Dean and Lewis (2005) studied the maximum loss in uniformly balanced repeated measurements designs (UBRMDs) in t periods when subjects may drop out after period t-m.
We will further study UBRMDs under the subject dropouts.
(1) We will derive the "best" UBRMDs for the situation where all subjects may drop out in the final period and provide methods for constructing these designs.
(2) We will study UBRMDs under subject dropout for the model where subject effects are random and show that the Low, Lewis and Prescott (1999) result on lack of connectedness of the Williams Latin Square of order 4 is no longer valid. Compound symmetry and AR (1) covariance structures will be considered.
(3) Expected loss under various dropout probabilities will be studied.
Oct. 29, 2008
Prof. Wei Biao Wu :
4:15 p.m. in SEO 612
Abstract
A popular framework for false discovery control is the random effects model in which the null hypotheses are assumed to be independent. I will generalize this random effects model to a conditional dependence model which allows dependence between null hypotheses. The dependence can be useful to characterize the spatial structure of the null hypotheses. Asymptotic properties of false discovery proportions and numbers of rejected hypotheses are explored and a large-sample distributional theory is obtained.
The talk is based on the paper: Wu, W. B. (2008) On false discovery control under dependence, Ann. Statist. 36, 364--380.
Nov. 5, 2008
Prof. Nanny Wermuth :
4:15 p.m. in SEO 612
Abstract
A joint density of several variables may satisfy a possibly large set of independence statements, called its independence structure. Often this structure is fully representable by a graph that consists of nodes representing variables and of edges that couple node pairs. We consider joint densities of this type, generated by a stepwise process in which all variables and dependences of interest are included. Otherwise, there are no constraints on the type of variables or on the form of the distribution generated. For densities that then result after marginalising and conditioning, we derive what we name the summary graph. It is seen to capture precisely the independence structure implied by the generating process, it identifies dependences which remain undistorted due to direct or indirect confounding and it alerts to possibly severe distortions of these two types in other parametrizations. We use operators for matrix representations of graphs to derive matrix results and translate these into to special types of path.
Nov. 12, 2008
Prof. Elton P Hsu :
4:15 p.m. in SEO 612
Abstract
How fast does a transient Brownian motion escape to infinity is an interesting question. For example, it is well known that Brownian motion in a euclidean
space of dimension 3 and higher escapes to infinity at approximately the rate of the square root of time. Which geometric quantity controls the rate of escape
is an interesting question. We will explain that the speed of Brownian motion can be effectively controlled by the volume growth rate of a complete Riemannian manifold. A precise integral
criterion is obtained which in many cases is sharp.
Nov. 19, 2008
Prof. Brad Efron :
4:15 p.m. in Lecture Center D2
Abstract
Classical prediction methods such as Fisher's linear discriminant
function were designed for small-scale problems, where the number N
of candidate predictors was much smaller than the number of observations
n. Modern scientific devices often reverse this situation. A micro-
array analysis, for example, might include n=100 subjects measured
on N=10,000 genes, each of which is a potential predictor. I will
discuss "Ebay", an empirical Bayes prediction algorithm designed to
handle N >> n situations. It is closely related to the Shrunken
Centroids algorithm of Tibshirani, Hastie, Narasimhan, and Chu.
Dec. 3, 2008
Prof. Min Zhang :
4:15 p.m. in SEO 612
Abstract
High throughput biotechnologies such as microarray and next-generation
sequencing permit simultaneous measurements of enormous bodies of
expression and sequence information. However, the number of biological
samples is much smaller compared to the number of available predictors.
Statistically, we are challenged by the large number of parameters but
small number of observations. To tackle this issue, we proposed a
two-step variable selection procedure to reduce the dimension in the
first stage where Gibbs sampler was developed to stochastically search
through low-dimensional subspaces. With reduced number of variables,
either Bayesian variable selection or traditional approaches can be
employed in the second stage. The methods are evaluated via simulation
studies and we also applied them to real data sets, including QTL mapping
data and gene expression data.
Jan. 21, 2009
Prof. Xiaofeng Shao :
4:15 p.m. in SEO 612
Abstract
This talk consists of two parts. In the first part, we will talk about testing for white noise and its applications to goodness-of-fit of long memory time series models. The limitation of the current asymptotic theory for portmanteau tests will be pointed out and new theoretical results will be discussed. In the second part, we will introduce generalized portmanteau type test statistics in the frequency domain to test independence between two stationary time series. Unlike the existing tests, each time series is allowed to possess short memory, long memory or anti-persistence. Under the null hypothesis of independence, the asymptotic null distributions of the proposed statistics are standard normal. The results from a simulation study will also be presented.
Jan. 28, 2009
Prof. Wei Biao Wu :
4:15 p.m. in SEO 612
Abstract
A popular framework for false discovery control is the random effects model in which the null hypotheses are assumed to be independent. I will generalize this random effects model to a conditional dependence model which allows dependence between null hypotheses. The dependence can be useful to characterize the spatial structure of the null hypotheses. Asymptotic properties of false discovery proportions and numbers of rejected hypotheses are explored and a large-sample distributional theory is obtained.
The talk is based on the paper: Wu, W. B. (2008) On false discovery control under dependence, Ann. Statist. 36, 364--380.
Feb. 4, 2009
Prof. George Karabatsos :
3 p.m. in SEO 636
Abstract
Often, causal inference is conducted on the basis of the randomized
experiment. However, in many settings, a randomized experiment is
infeasible, because treatments cannot be directly assigned to subjects.
Specifically, it may not be possible for the investigator to assign
treatments to subjects, because of ethical concerns, or because of
excessive expense in terms of time or money. In such settings, causal
inference needs to be undertaken in an observational study, where the
subjects received different treatments, but the investigator did not
assign the treatments, and therefore the treatment assignment
probabilities are unknown.
Typically, in the practice of causal inference from observational studies,
a parametric model is assumed for the joint population density of
potential outcomes and treatment assignments, and possibly this is
accompanied by the assumption of no hidden bias. However, both assumptions
are questionable for real data, the accuracy of causal inference is
compromised when the data violates either assumption, and the parametric
assumption precludes capturing a more general range of density shapes
(e.g., heavier tail behavior and possible multi-modalities in the joint
density). We introduce a flexible, Bayesian nonparametric causal model to
provide more accurate causal inferences. The model makes use of a
stick-breaking prior distribution, which has the flexibility to capture
any multi-modalities, skewness and heavier tail behavior in this joint
population density, while accounting for hidden bias. We prove the
asymptotic consistency of the posterior distribution of the model. Also,
we illustrate our Bayesian nonparametric causal model through the analysis
of small genetic data set, and a large data set of Chicago public schools.
Feb. 11, 2009
Prof. Abhyuday Mandal :
4:15 p.m. in SEO 612
Abstract
Functional magnetic resonance imaging (fMRI) is considered one of the leading technologies for studying human brain activity in response to mental stimuli.
With sophisticated allocations of stimuli, researchers can gather valuable fMRI time series and acquire precise information about human brain activity.
However, due to the nature of fMRI experiments, the underlying design space is very large and irregular. This makes it difficult to find an optimal design
that simultaneously accomplishes various goals of a study and fulfills the scientific restrictions. Here we propose an efficient approach to find optimal
experimental designs for event-related functional magnetic resonance imaging (ER-fMRI). We consider multiple objectives, including estimating the
hemodynamic response function (HRF), detecting activation, circumventing psychological confounds and fulfilling customized requirements.
Taking into account these goals, we formulate a family of multi-objective design criteria and develop a genetic-algorithm-based technique
to search for optimal designs. Our proposed technique incorporates existing knowledge about the performance of fMRI designs, and its usefulness
is shown through simulations. We also find designs yielding higher estimation efficiencies than m-sequences. When the underlying model
is with white noise and a constant nuisance parameter, the stimulus frequencies of the designs we obtained are in good agreement with the optimal
stimulus frequencies derived by Liu and Frank, 2004, NeuroImage 21, 387-400. In terms of CPU time and achieved design efficiency, we demonstrate
that our approach outperforms the methodologies known hitherto. (Joint research with Ming-Hung (Jason) Kao, John Stufken and Nicole Lazar)
Feb. 13, 2009
Dr. Arlene Ash :
2 p.m. in SEO 612
Abstract
US elections are very complicated, providing many opportunities for
inadvertent and malicious errors. I will examine the statistical evidence
regarding a "failed election" in Florida's 13th Congressional District in
2006 and discuss some lessons learned. I will discuss additional ways in
which elections can be problematic and how statistics can be used to
identify problems and explore solutions. Finally, I will describe work in
progress with state election officials to improve the efficiency and
effectiveness of post-election audits.
Feb. 25, 2009
Sonja Petrovic :
4:15 p.m. in SEO 612
Abstract
Algebraic statistics is a maturing discipline whose main focus is the study of statistical models using the tools from algebraic geometry and computational algebra. The main concept is that statistical models are algebraic varieties. Algebraic approach can be used to provide Markov bases for the models, or to compute the maximum likelihood degree. Some of the best studied models so far are contingency tables, conditional independence models and graphical models including latent class.
Algebraic statistics has also found applications in computational biology and phylogenetics. Some recent work shows that it can be used as a powerful tool for phylogenetic tree reconstruction and for model identifiability problems.
This will be an introductory talk to explain some of the main concepts of the field, illustrated on a few examples.
March 4, 2009
Prof. Liping Tong :
3 p.m. in SEO 636
Abstract
Accurate characterization of haplotype structure and diversity is a key challenge in statistical genetics. Attempts to apply findings from genome wide association studies to populations not included in the discovery phase present unique challenges in terms of the statistical methods. In this talk, I propose a new statistic to assess and compare the haplotype variations among populations which is particularly suited to this emerging challenge. Subsequently I show that this statistic follow a weighted chi-square distribution and how to use a chi-square distribution to approximate it. This approximation is very important since no other haplotype similarity tests have (correctly) used approximate theoretical distributions. In stead, the computational intensive permutation tests are generally performed, which limit the application of haplotype-based comparisons to the whole genome wide studies. In the simulation studies, I first discuss the performance of the approximate distribution under different definitions of similarity matrix, and then compare the power of my new method with the ones proposed by others. At last, this method is applied to the HapMap data to test population differences based on haplotypes on chromosome 2 in the region surrounding the LCT gene (135.3-136.9 Mb).
March 18, 2009
Prof. Fang Li :
4 p.m. in SEO 612
Abstract
In this talk, we introduce the Multivariate Theil-Sen Estimators (MTSE) in a multiple linear regression model generalizing the TSE in a simple linear model. We demonstrate that the proposed estimators are robust with bounded influence function. We also show that the MTSE is consistent under very mild conditions. When the covariates are independent and identically distributed, the related criterion statistics defining the MTSE are the standard U-statistics. However when the covariates are deterministic, the related criterion statistics are no longer U-statistics. We then use the Convexity Lemma of Pollard to prove its asymptotic normality under mild condition. This method can also be easily generalized to the case when the covariates are independent and identically distributed.
April 8, 2009
Prof. Tim McMurry :
4:15 p.m. in SEO 612
Abstract
I will discuss the usefulness of infinite order kernels in nonparametric smoothing problems. Particular focus will be paid to nonparametric regression, but the ideas apply to many related problems, including density and spectral density estimation. It will be shown that estimators using infinite order kernels are completely and automatically adaptive to the underlying function being estimated, produce favorable results in simulation, and substantially facilitate construction of confidence intervals.
April 15, 2009
Prof. Hua Liang :
4:15 p.m. in SEO 636
Abstract
We investigated two semi-parametric models, partially linear model and generalized partially linear model, with error-prone covariates. A
correction-for-attenuation method was developed for estimating the parameter of interest in the partially linear model. The resulting
estimator was shown to be consistent and its asymptotic distribution theory has been derived. Consistent standard error estimates using
sandwich-type ideas were also developed. For generalized partially linear model, we proposed estimators of parameter and nonparametric
function by using local linear regression, simulation extrapolation technique, and generalized estimating equation. The asymptotic normality
of the estimators of the parameter, the bias and variance of the estimators of the nonparametric component were derived under appropriate
assumptions. We illustrated the numerical performance of the proposed methods via simulation and examples, discussed the potential topics for
further work.
April 22, 2009
Prof. Min Zhang :
4:15 p.m. in SEO 612
Abstract
High throughput biotechnologies such as microarray and next-generation sequencing permit simultaneous measurements of enormous bodies of expression and sequence information. However, the number of biological samples is much smaller compared to the number of available predictors. Statistically, we are challenged by the large number of parameters but small number of observations. To tackle this issue, we proposed a two-step variable selection procedure to reduce the dimension in the first stage where Gibbs sampler was developed to stochastically search through low-dimensional subspaces. With reduced number of variables, either Bayesian variable selection or traditional approaches can be employed in the second stage. The methods are evaluated via simulation studies and we also applied them to real data sets, including QTL mapping data and gene expression data.
April 29, 2009
Wei Zheng, PhD candidate :
4:15 p.m. in SEO 612
Abstract
The statistical optimality and efficiency of crossover designs for the purpose of comparing several test treatments with a control treatment when the subject effects are random depend heavily on the unknown ratio theta of the variance of subject effects and the error variance. However, it is proved that if the class of competing designs contains a totally balanced test-control incomplete crossover designs (TBTCI), as defined by Hedayat and Yang (2005), then this TBTCI design is simultaneously A- and MV-optimal for all values of theta. This result is essentially a generalization of a result in Hedayat and Yang (2005) since their statistical model is based on fixed subject effects, where the Fisher information matrix would be identical to that of random subject effect model when theta goes to infinity. Partial works on the construction of the designs are carried out.
Sept. 2, 2009
Prof. Lan Xue :
3 p.m. in SEO 636
Abstract
We propose a penalized polynomial spline method for
simultaneous model estimation and variable selection in additive models. It
approximates nonparametric functions by polynomial splines, and
minimizes the sum of squared errors subject to an additive penalty on norms
of spline functions. This approach sets estimators of certain function
components to zero, thus performing variable selection. Under mild
conditions, we show that the newly proposed method estimates the non-zero
function components in the model with the same optimal mean square
convergence rate as the standard polynomial spline estimators, and correctly
sets the zero function components to zero with probability approaching one,
as $n$ goes to infinity. Besides being theoretically justified, the proposed
method is easy to understand and straightforward to implement. Extensive
Monte Carlo simulation studies show the newly proposed method compares
favorably with the existing ones in finite sample performance. We also
illustrate the use of the proposed method by analyzing two data sets.
Sept. 9, 2009
Prof. Dabao Zhang :
3 p.m. in SEO 636
Abstract
We propose a penalized orthogonal-components regression
(POCRE) for large p small n data. Orthogonal components are sequentially
constructed to maximize, upon standardization, their correlation to the
re-
sponse residuals. A new penalization framework, implemented via empiri-
cal Bayes thresholding, is presented to effectively identify sparse
predictors
of each component. POCRE is computationally efficient owing to its se-
quential construction of leading sparse principal components. In
addition,
such construction offers other properties such as grouping highly
correlated
predictors and allowing for collinear or nearly collinear predictors.
With
multivariate responses, POCRE can construct common components and
thus build up latent-variable models for large p small n data. This is
an joint work with Yanzhu Lin and Min Zhang.
Sept. 16, 2009
Prof. Ahmad Reza Soltani :
3 p.m. in SEO 636
Abstract
An approach for proving the strong consistency of certain estimators for unknown parameters in the context of statistical inference is given. This approach is based on an application of the Implicit Function Theorem in Hilbert spaces, and can be applied to the random samples consisting of univariate, multivariate or infinite dimensional random elements.
Sept. 23, 2009
Prof. Guang Cheng :
3 p.m. in SEO 636
Abstract
Consider M-estimation in a semiparametric model that is characterized by a
Euclidean parameter of interest and a nuisance function parameter. We show
that, under general conditions, the bootstrap is asymptotically consistent
in estimating the distribution of the M-estimate of Euclidean parameter;
this is, the bootstrap distribution asymptotically imitates the distribution
of the M-estimate. We also show that the bootstrap confidence set has the
asymptotically correct coverage probability. These general conclusions hold,
in particular, when the nuisance parameter is not estimable at root-n rate.
Our results provide a theoretical justification for the use of bootstrap as
an inference tool in semiparametric modelling and apply to a broad class of
bootstrap methods with exchangeable bootstrap weights. A by-product of our
theoretical development is the second order asymptotic linear expansion of
the (bootstrap) M-estimate. Joint work with Jianhua Huang at Texas A&M University.
Sept. 30, 2009
Prof. Zhengjun Zhang :
3 p.m. in SEO 636
Abstract
Various correlation measures have been introduced in statistical inferences and applications. Each of them may be used in measuring association strength of the relationship, or testing independence, between two random variables. \textsl{The quotient correlation} is defined here as an alternative to Pearson's correlation that is more intuitive and flexible in cases where the tail behavior of data is important. It measures nonlinear dependence where the regular correlation coefficient is generally not applicable. One of its most useful features is a test statistic that has high power when testing nonlinear dependence in cases where the Fisher's $Z$-transformation test may fail to reach a right conclusion. Unlike most asymptotic test statistics, which are either normal or $\chi2$, this test statistic has a limiting gamma distribution (henceforth \textsl{the gamma test statistic}). More than the common usages of correlation, the quotient correlation can easily and intuitively be adjusted to values at tails. This adjustment generates two new concepts -- the tail quotient correlation and the tail independence test statistics, which are also gamma statistics. Due to the fact that there is no analogue of the correlation coefficient in extreme value theory, and there does not exist an efficient tail independence test statistic, these two new concepts may open up a new field of study. In addition, an alternative to Spearman's rank correlation: a rank based quotient correlation is also defined. The advantages of using these new concepts are illustrated with simulated data, and real data analysis of internet traffic, tobacco markets, financial markets...
Oct. 7, 2009
Prof. Hedibert Lopes :
3 p.m. in SEO 636
Abstract
This paper develops efficient sequential learning methods for the estimation of general mixture models, by working directly with particles based on conditional sufficient information. We provide an alternative to existing inference techniques that will be especially relevant in on-line estimation settings and for large,high-dimensional data-sets. With each new observation, particles are updated in two steps: first, resampling with weights proportional to the implied predictive probability distribution and, secondly, propagating the next latent mixture allocation and implicitly sampling the next particle vector. We introduce the methodology in the context of finite mixture models, before extending to any nonparametric mixture model with an available predictive probability function and focusing, in particular, on Dirichlet Process mixture models. In addition, we show that the algorithm provides a natural estimate for sequential Bayes factors and can facilitate selection between competing dynamic models. The framework is illustrated with numerous real and simulated data examples. (This is joint work with Carlos Carvalho, Nicholas Polson and Matt Taddy)
Oct. 28, 2009
Prof. Jiping Wang :
3 p.m. in SEO 636
Abstract
Suppose D distinct species are observed from an infinite population consisting of N (unknown) distinct species. The estimate of N from popular nonparametric methods can be substantially biased downward, while parametric
approaches assuming a smooth abundance curve in general lack robustness. In this paper we propose a Poisson-compound Gamma approach, where the species
abundance distribution Q is modeled as a Gamma mixture. We first show a nesting property of the Gamma mixture model, under which an arbitrary finite
Gamma mixture can be uniquely re-written as a new Gamma mixture with components sharing a unified shape parameter, and mixed in the mean parameter. Thereby Q can be estimated using nonparametric maximum likelihood
method for any given a. We further propose a least-squares cross-validation procedure for choice of the shape parameter to attain the desired smoothness
of Q while controlling the goodness of fit of the model. The competitive performance of the resulting N-estimator is demonstrated using numerical studies and newly arising genomic data.
Nov. 4, 2009
Prof. Peter Qian :
3 p.m. in SEO 636
Abstract
We introduce a new type of design, called nested Latin hypercube design, for sequential integration and multi-fidelity computer modeling. A nested Latin hypercube design is defined to be a special Latin hypercube design that contains a smaller Latin hypercube design as a subset. Such designs are constructed by exploiting nested structures in random permutations. The constructed designs are also useful for solving stochastic optimization problems, including stochastic programs, the Monte Carlo EM algorithm and chance-constraint problems.
Nov. 11, 2009
Prof. Hanxiang Peng :
3 p.m. in SEO 636
Abstract
In this talk, we discuss the asymptotic properties
of a semiparametric multiplicative hazard model when the
relative risk is expressed as a first order continuously
differentiable parametric function. We show that the log-
the partial likelihood function of the model is locally
concave for an arbitrary continuously differentiable relative
risk under suitable conditions. Then we derive the
existence and uniqueness of the MPLE and show consistency.
Using the convexity lemma and characterization of minimizers,
we demonstrate that the MPLE of the parameter is asymptotically
normal. As an application, we exhibit that the MPLE of the
parameter in a model in which the log- the relative risk is
expressed as a free-knot spline with knots in covariates uniquely
exists in a neighborhood of the true parameter value and is
consistent and asymptotically normal. In particular, we derive the
asymptotic normality of the MPLE of the parameter in a model in
which the log- relative risk is expressed as a free-knot
quadratic spline which has first order continuous derivative.
Nov. 18, 2009
Prof. Dulal Bhaumik :
3 p.m. in SEO 636
Abstract
A new statistical methodology is developed for analysis of spontaneous adverse event reports from post-marketing
drug surveillance data. The method involves both empirical Bayes and fully-Bayes estimation of rate multipliers
for each drug within a class of drugs, for a particular adverse event, based on a mixed-effects Poisson regression model.
Both parametric and semi-parametric models for the random effect distribution are examined. The method is applied to
data from FDA`s Adverse Event Reporting System (AERS) on the relationship between antidepressants and suicide.
We obtain point estimates and 95% confidence intervals for the rate multiplier for each drug (e.g., antidepressants),
which can be used to determine if a particular drug has an increased risk of association with a particular adverse event
(e.g., suicide). Confidence intervals that do not include 1.0 provide evidence for either significant protective or
harmful associations of the drug and the adverse effect. We also examine empirical Bayes, parametric Bayes and
semi-parametric Bayes estimators of the rate multipliers and associated confidence intervals. Results of our analysis
of the FDA AERS data revealed that newer antidepressants are associated with lower rates of suicide. This finding
contradicts previous findings of FDA that newer antidepressants are causally related to increased suicidal thinking
in children and young adults. Finally, we suggest changes in the AERS system to improve our ability to discover these
adverse events.
Jan. 15, 2010
Wei Zheng :
2 p.m. in SEO 612
Abstract
In crossover designs, it suffices to consider a linear model with effects of subjects, periods, direct and first-order carryover effects of treatments in most applications. We allowed the subject effects to be random, and thus included the fixed subject effects model as a special case. Efficient designs are proposed under this more flexible model. We also studied statistical properties and methods of constructions of these designs.
For the remaining few minutes, I will talk about asymptotic behavior of sample autocovariances of long-memory linear processes. Also, some of my future interests will be addressed at the end.
Jan. 27, 2010
Ying Zhou :
3 p.m. in SEO 636
Abstract
D-optimal designs for nonlinear chemical kinetics model and 2n-compartment
models are investigated. We develop a method to obtain lower bounds of the
maximal numbers of D-optimal design points to facility the computational
search for the optimal design points in practice. We also investigate the
conditions when the D-optimal design is the saturated design for bounded and
unbounded design spaces. For each model discussed, the D-efficiency when the
parameter misspecification happens is discussed.
Feb. 3, 2010
Cuilan Zhang :
3 p.m. in SEO 636
Abstract
In clinical trials, it is important to find optimal allocation strategies to control the trials and benefit the participants. In our paper, we obtained explicit optimal allocation strategies for multiple treatments(2 and 3 treatments), in order to minimize expected total number of responses larger than a threshold, while guaranteeing a minimum power for the hypothesis that all treatments follow the same distribution. The assumed distributions are the standard Weibull models with common shape parameter. This optimal allocation strategy can be implemented by Doubly- biased coin design sequentially.
Feb. 17, 2010
Han Xiao :
3 p.m. in SEO 636
Abstract
Let $X^{(n)}=(X_{ij})$ be a $p \times n$ data matrix, where the $n$ columns form a random sample of size $n$ from a certain $p$-dimensional
distribution. Let $R^{(n)}=(\rho_{ij})$ be the $p \times p$ sample correlation coefficient matrix of $X^{(n)}$; and $S^{(n)} =
(1/n)X^{(n)}\left(X^{(n)}\right)^{\ast}-\bar{X}\bar{X}^{\ast}$ be the sample covariance matrix of $X^{(n)}$, where $\bar{X}$ is the mean vector of the
$n$ observations. Assuming that $X_{ij}$'s are independent and identically distributed with finite fourth moment, we show that the smallest eigenvalue
of $R^{(n)}$ converges almost surely to the limit $(1-\sqrt{c}\,)^2$ as $n \rightarrow \infty$ and $p/n \rightarrow c \in (0\,,\,\infty)$. We
accomplish this by showing that the smallest eigenvalue of $S^{(n)}$ converges almost surely to $(1-\sqrt{c}\,)^2$.
Feb. 24, 2010
Yuqing Tang :
3 p.m. in SEO 636
Abstract
Evaluating individual agreement between different raters are often of interest
in method comparison studies and in reliability studies. We propose a general
comparison model and create a
exible setting such that any subset of raters can
be selected as test or reference raters and the individual agreement between them
can be assessed. Two comparative agreement indices, the individual difference ratio
(IDR) and within difference ratio (WDR), are proposed. IDR is a non-inferiority
assessment such that the individual reading from different raters can not be inferior
to the replicated readings within the same rater. When there is one test rater and
one reference rater, our approach degenerates to FDA's method for evaluating
individual bioequivalence under relative scale. WDR is a superiority assessment
such that the precision of selected test raters can be better than that of selected
reference raters. GEE approach is used for estimation and inference. Simulation
study is conducted to assess the performance of our approach and the result shows
that our method works well for both continuous data and categorical data.
March 3, 2010
Prof. David Degras :
3 p.m. in SEO 636
Abstract
This work deals with the construction of simultaneous confidence bands (SCB) in functional mixed-effects models. SCB are useful graphical and analytic tools in data exploration, model construction, estimation of fixed and random effects, covariance estimation, prediction, and inference. The model under study can handle dummy variables and functional predictors, and its nonparametric covariance structure accounts for dependence in the data in a more flexible way than other models based on smoothing spline or wavelet representations. The SCB-based method allows to test local and nonparametric alternatives on the model components as opposed to the usual F-type tests. Some asymptotic theory is derived and a numerical study is presented to compare our method with other approaches currently in practice. We also illustrate our methodology by applying it to a real data set.
March 17, 2010
Prof. Annie Qu :
3 p.m. in SEO 636
Abstract
Model selection of correlation structure is a challenging problem because it
involves a higher order of moments than model selection of covariates only.
In addition, the high dimension of the correlation parameters could make the
estimation of those parameters unreliable since the number of repeated
measurements might be relatively small compared to the dimension of the
correlation parameters. However, the correct specification of the correlation
structure plays an important role in improving estimation efficiency for
clustered data. We propose to select the correlation structure for clustered data
from a number of candidate structures through a group-wise basis matrices
selection strategy. The proposed method has the advantages of not requiring
the likelihood function and of being computationally efficient. Also, the method
can identify complex correlation structures. Furthermore, it is applicable for
both continuous and discrete response data. In theory, we show that the
proposed method enjoys the oracle property of selecting the true correlation
structure consistently and estimating the correlation parameters with the same
asymptotic normal distribution as if the true structure is known. This is joint
work with Jianhui Zhou of University of Virginia.
March 31, 2010
Prof. Fangfang Wang :
3 p.m. in SEO 636
Abstract
We propose a general GARCH framework that allows the use of different
frequency returns to model conditional heteroskedasticity. We call the class of models High
FrequencY Data-Based PRojectIon-Driven GARCH models as the GARCH dynamics are driven by
what we call HYBRID processes. We study three broad classes of HYBRID processes: (1)
parameter-free processes that are purely data-driven, (2) structural HYBRIDs where one
assumes an underlying DGP for the high frequency data and finally (3) HYBRID filter
processes. We develop the asymptotic theory of various estimators and study their
properties in small samples via simulations.
This is joint work with Eric Ghysels (University of North Carolina at Chapel
Hill) and Xilong Chen (SAS Institute Inc.).
April 7, 2010
Prof. Sijian Wang :
3 p.m. in SEO 636
Abstract
The linear mixed effects model (LMM) is widely used in the analysis of clustered or longitudinal data. In the practice of LMM, inference on the structure of random effects component is of great importance not only to yield proper interpretation of subject-specific effects but also to draw valid statistical conclusions. This task of inference becomes significantly challenging when a large number of fixed effects and random effects are involved in the analysis. The difficulty of variable selection arises from the need of simultaneously regularizing both mean model and covariance structures, with possible parameter constraints between the two. In this paper, we propose a novel method of regularized restricted maximum likelihood to select fixed and random effects simultaneously in the LMM. The Cholesky decomposition is invoked to ensure the positive-definiteness of the selected covariance matrix of random effects, and selected random effects are invariant with respect to the ordering of predictors appearing in the model. We develop a new algorithm that solves the related optimization problem effectively, in which the computational load turns out to be comparable with that of the Newton-Raphson algorithm for MLE or REML in the LMM. We also investigate large sample properties for the proposed estimation, including the oracle property. Both simulation studies and data analysis are included for illustration. This is a joint work with Peter XK Song and Ji Zhu.
April 9, 2010
Prof. Abhyuday Mandal :
2 p.m. in SEO 612
Abstract
Functional magnetic resonance imaging (fMRI) is an important tool for
scientists studying brain function. FMRI data are complex in nature: they
are massive in size and a low signal-to-noise level makes the elimination of
some noise prior to model fitting desirable for improved identification of
true brain activity. We propose two methods of reducing this noise:
generalized indicator functional analysis and a hidden Markov model. Brain
regions showing increased fMRI signal while subjects engaged in a
visual/spatial motor task are identified using concepts from social network
analysis and statistical mechanics. Conditional probabilities of activation
given the degree to which pairs of voxels are related are modeled for three
groups: people with schizophrenia, their asymptomatic relatives, and control
subjects. We compare the conditional probability maps obtained for each
group to evaluate for between-group differences in extent of task-related
signal.
(Joint research with Ana M. Bargo, Lynne Seymour, Jennifer McDowelly, and
Nicole A. Lazar)
April 14, 2010
Prof. Noelle Samia :
3 p.m. in SEO 636
Abstract
The open-loop Threshold Model proposed by Tong (1990) is a stochastic piecewise-linear
regression model useful for modeling conditionally normal response time-series data. How-
ever, in many applications, the response variable is conditionally non-normal, e.g. Poisson
or binomially distributed. We generalize the open-loop Threshold Model by introducing
the Generalized Threshold Model (GTM). Specifically, it is assumed that the conditional
probability distribution of the response variable belongs to the exponential family, and the
conditional mean response is linked to some piecewise-linear stochastic regression function.
The consistency and limiting distribution of the maximum likelihood estimator are derived.
We illustrate the GTM with a real application on the annual number of human bubonic
plague cases in Kazakhstan. This is based on joint work with Professor Kung-Sik Chan
(University of Iowa) and Professor Nils C. Stenseth (University of Oslo).
April 21, 2010
Prof. Hongmei Jiang :
3 p.m. in SEO 636
Abstract
Copy number changes, either amplification or deletion of DNA
materials, have been linked with cancer and other diseases.
High-throughput DNA arrays enable simultaneous measurements of copy
number on the genome-wide level. Given the intensity measurements from a
single sample, methodologies and algorithms based on diverse techniques
such as Hidden Markov Models, binary segmentation, and mixture models,
have been developed to divide the genome into segments or regions of
equal copy number. In this talk we will discuss how to find regions of
recurrent gains or losses of DNA fragments across multiple cancer
samples; how to identify chromosome regions which are associated with
clinical variables such as relapse and survival outcome. The methods
will be demonstrated on a liver cancer data set.
April 23, 2010
Prof. John Morgan :
4:15 p.m. in SEO 636
Abstract
Standard optimality arguments for designed experiments rest on the
assumption that all treatments are of equal interest. A notable
exception is found in the ``test treatment versus control" (TvC)
literature, where the control is allocated special status.
Optimality work there has focused on all pairwise comparisons with
the control, making no explicit account of how well test treatments
are compared to one another. If the latter are also of consequence,
it would be preferable to choose a design reflecting the relative
importance placed on contrasts involving the control to that placed
on contrasts of test treatments only. This talk develops the
\emph{weighted} optimality approach for situations such as this, so
that design selection may better reflect experimenter goals.
When evaluating designs for comparing $v$ treatments, the basic idea
is to assign weights $w_1,\ldots,w_v$ ($\sum_iw_i=1$) to account for
differential treatment interest. In experiments with a control, and
equal interest in the test treatments, this means weight $w_1$ is
assigned to the control, and weight $w_2=(1-w_1)/(v-1)$ to each test
treatment. These weights enter the evaluation through optimality
measures, leading to, for example, weighted versions of the popular
A, E, and MV measures of design efficacy. Families of
weighted-optimal designs are identified under these criteria.
Compared to their unweighted versions, it is shown that they less
frequently agree on the best design. The classical approach in TvC
design is shown to be a limiting case of the theory developed here.
April 28, 2010
Prof. Bruce D. Spencer :
3 p.m. in SEO 636
Abstract
Criminal trials may be viewed as complex classification procedures where the verdict represents classification as guilty or not guilty. Assessing the accuracy of verdicts is difficult because the "true" state of the defendant typically is unknown, and those cases where it is known are atypical. Yet, average accuracy of verdicts in criminal cases can be studied systematically and empirically provided we can obtain a second (or even a third) rating of the verdict. For example, in a jury trial the judge can also be asked for a verdict, as in the National Center for State Courts (NCSC) study of criminal cases from four jurisdictions in 2000-01. That study, like the famous Kalven-Zeisel study of the 1950s, showed only modest agreement between the judge and jury. Estimates of overall accuracy of verdicts are easily developed from the judge-jury agreement rate, and under plausible conditions the estimates of accuracy are optimistic. Estimates of false conviction rates and false acquittal rates, are more challenging, and are developed for the NCSC data with the use of log-linear latent class models. Those models, as well as models based on more than two raters, depend on stronger assumptions than the estimates of overall accuracy based on agreement rates. Numerical estimates of verdict accuracy are presented for the NCSC data and sources of uncertainty in the estimates are discussed, with particular attention to the effect of invalidity of the latent class. The estimates of the false conviction rates and false acquittal rates lead to questions about the appropriate balance of errors. Limitations of statistical decision theory for finding an optimal balance will be discussed.
Aug. 25, 2010
Junhui Wang :
3 p.m. in SEO 636
Sept. 8, 2010
Lingsong Zhang :
3 p.m. in SEO 636
Abstract
Distance Weighted Discrimination (DWD) has recently been proposed as an
attractive classification method. In this paper, we first show Fisher
consistency of the DWD method, which justifies its use when there are
sufficient data. However, the DWD classifier is not sparse, which makes
the interpretation and prediction performance less attractive. We
propose several sparse DWD methods, which incorporate variable selection
techniques in classification using penalized loss functions to estimate
the true hyperplane. We show that when an appropriate penalty is used,
the sparse DWD method is consistent and the estimated normal vector has
the oracle property under suitable conditions. We evaluate the finite
sample performance of the proposed methods using simulations and
illustrate the methods with an application to the Faroe island proteomic
biomarker data.
Sept. 15, 2010
Sayan Mukherjee :
3 p.m. in SEO 636
Abstract
The focus is on the problem of supervised dimension reduction (SDR). We
first formulate the problem with respect to the inference of a geometric
property of the data, the gradient of the regression function with respect
to the manifold that supports the marginal distribution. We provide an
estimation algorithm, prove consistency, and explain why the gradient is
salient for dimension reduction. We then reformulate SDR in a
probabilistic framework and propose a Bayesian model, a mixture of inverse
regressions. In this modeling framework the Grassman manifold plays a
prominent role.
Sept. 22, 2010
Hongmei Liu :
3 p.m. in SEO 636
Abstract
Basic ideas of survey sampling were utilized in a project for STAT431 "Introduction to Survey Sampling" at the University of Illinois at Chicago, taught and supervised by Professor Hedayat in the fall 2009, to
estimate the mis-shelving rate of books at the University of Illinois-Chicago (UIC) Daley library. We use this project as an example to illustrate how to design a sampling plan, sample size, and how to estimate
rates at which books are mis-shelved.
Sept. 29, 2010
Timothy E. O'Brien :
3 p.m. in SEO 636
Abstract
Researchers often find that nonlinear regression models are more applicable for
modelling various biological, physical and chemical processes than are linear ones since
they tend to fit the data well and since these models (and model parameters) are more
scientifically meaningful. These researchers are thus often in a position of requiring
optimal or near-optimal designs for a given nonlinear model. A common shortcoming of
most optimal designs for nonlinear models used in practical settings, however, is that
these designs typically focus only on (first-order) parameter variance or predicted
variance, and thus ignore the inherent nonlinear of the assumed model function.
Another shortcoming of optimal designs is that they often have only p support points,
where p is the number of model parameters.
Measures of marginal curvature, first introduced in Clarke (1987) and further developed
in Haines et al (2004), provide a useful means of assessing this nonlinearity. Other
relevant developments are the second-order volume design criterion introduced in
Hamilton and Watts (1985) and extended in O'Brien (1992, 2010), and the second-order
MSE criterion developed and illustrated in Clarke and Haines (1995).
This talk examines various robust design criteria and those based on second-order
(curvature) considerations. These techniques, coded in the GAUSS and SAS/IML
software packages, are illustrated with several examples including one from a preclinical
dose-response setting encountered in a recent consulting session.
Oct. 6, 2010
Yue Yu :
3 p.m. in SEO 636
Abstract
Sliced Inverse Regression (SIR) proposed by Ker-Chau Li (1991) is a widely
used semiparametric technique to reduce the dimensions of regression
problems. But the microeconomics data are time dependent and usually highly
correlated. In our study, we use clustering methods along with SIR, in order
to reduce the multicollinearity and the difficulty of choosing the
dimensions. And the dynamic version of cluster SIR methods is used to
analyze the autoregressive model of the microeconomics data. Our simulation
result shows that the dynamic cluster SIR is superior, comparing with the
empirical accuracy of all the models in Stock and Watson's paper (2005) for
forecasting U.S. macroeconomic time series over a 30-year period.
Oct. 13, 2010
Yuan Xu :
3 p.m. in SEO 636
Abstract
We develop a joint model of longitudinally observed cognitive data and survival data to the onset of dementia. We incorporate latent random change points in the model representing an accelerated cognitive decline prior to the onset of dementia. We aim to investigate how different covariates of subjects, such as baseline age, education and genetic risk factors, affect the timing of cognitive decline acceleration. We also assess how different groups of subjects behave on cognitive decline before and after the change point. The model combines a longitudinal mixed effects model with a Cox proportional hazards model connected by a random change point with a log normal distribution. The parameters are estimated by the maximum likelihood method through an ECM algorithm. Compared with joint models with change points developed previously by other authors, our model has several advantages. First, our model uses the semi-parametric Cox model instead of a parametric model for the survival data, therefore is more flexible to different survival distributions. Second, we use the maximum likelihood method and an ECM algorithm to estimate the parameters to avoid the prior assumptions on model parameters. Third, we propose a compromised Fisher information method other than profile likelihood method to obtain a better estimation of the standard errors of the MLEs for model parameters. Finally, the proposed model is successfully implemented to study the preclinical acceleration on the rate of cognitive decline as well as its implication on the risk of developing dementia.
Oct. 20, 2010
Troy Hernandez :
3 p.m. in SEO 636
Abstract
Bus arrival prediction times are useful to many passengers of public
transportation. There are many features of the environment that can be
used to predict bus arrival times and many possible representations of
these features.
Two representations used to predict bus arrival times are discussed: a
memoryless representation that uses only the current state of the bus
and a full-memory or trajectory-based representation that uses the
full history of the bus run.
Oct. 27, 2010
Jun Xie :
3 p.m. in SEO 636
Abstract
Many statistical classification methods, e.g., Fisher's linear
discriminant analysis, cannot be directly applied to high dimensional
data, where the number of variables is larger than the sample size. While
high dimensional data analysis has been broadly discussed in statistics
community, the impact of dimensionality on classifications is poorly
understood. We examine and compare high dimensional classification
methods through an application in pharmacogenomics research, where
high-dimensional gene expression microarray data are used to predict
patients' responses to a drug. Compared with most gene expression
classification studies to detect strong signals, for instance tumor
versus normal, a classifier between patients' response and non-response
is more challenging and may be nonlinear. We introduce several new
classification methods, including a sparse linear discriminant method,
random projection, and a distribution based classification involving
second-order interactions, as potential tools to deal with high
dimensionality. We also want to call attentions to theories of high
dimensional classification, where there are only few results available.
Nov. 3, 2010
Hongyuan Cao :
3 p.m. in SEO 636
Abstract
High-throughput screening has become an important mainstay for con-
temporary biomedical research. A standard approach is to get p-values and
adjust for multiple comparison in a manner that controls false discovery rate
(FDR). The concavity of p-value distribution under the alternative has been
a standard condition for developing many FDR procedures: Storey (2003),
Genovese and Wasserman (2004), Kosorok and Ma (2007). A more general
concept is the monotone likelihood ratio condition (MLRC) introduced in
Sun and Cai (2007). We show in this paper that the concavity assumption
can be violated for (i) a simple heteroscedastic normal mixture model and
(ii) dependent tests. Some interesting implications, including different testing procedures (step-up vs step-down), the choice of test statistic and the
power definition in multiple testing are discussed. This is joint work with
Wenguang Sun and Michael R. Kosorok.
Nov. 10, 2010
Elizabeth Gross :
3 p.m. in SEO 636
Abstract
Maximum likelihood estimation is a common problem explored in Algebraic Statistics, a field that focuses on the applications of algebraic geometry to the study of statistical models. In maximum likelihood estimation if the likelihood equations are algebraic, the maximum likelihood degree (ML degree) is a measure of the algebraic complexity of the estimation problem. It is the degree of the variety characterized by the system of likelihood equations, or, equivalently, the number of complex solutions of the system for generic data. The ML degree is specific to the statistical model and there are only two other classes of models for which an explicit formula for the ML degree is known. In this talk, we will look at the analysis of variance model with random effects and give an explicit formula for the ML degree. We also explore the number of feasible (real, positive) solutions computationally. This is joint work with Mathias Drton and Sonja Petrovic.
Nov. 17, 2010
Dibyen Majumdar :
3 p.m. in SEO 636
Abstract
Crossover studies are used in different areas of statistical applications and there is a substantial literature that focus on identifying and constructing efficient designs for these studies. However, even well-designed crossover studies often lose their statistical properties if subjects drop out before the end of the study. We will explore the problem of subject dropout and the effect on properties of the design, and search for efficient designs that are robust to subject dropout.
Nov. 23, 2010
Li Wang :
1:30 p.m. in SEO 636
Abstract
We study a class of generalized additive partial linear
models. We propose the use of polynomial spline smoothing for estimation
of nonparametric functions, and derive the quasi-likelihood based
estimators for the linear parameters. We establish asymptotic normality
for the estimators of the parametric components. The procedure avoids
solving big system of equations as in kernel-based procedures and thus
results in gains in computational simplicity. We further develop a class
of variable selection procedures for the linear parameters by employing
a nonconcave penalized likelihood, which is shown to have an oracle
property. Monte Carlo simulations and an analysis of a dataset from Pima
Indian diabetes study are presented for illustration.
Dec. 1, 2010
Yinxiao Huang :
3 p.m. in SEO 636
Abstract
In this paper, we give a unified analysis of both the nonparametric kernel
density estimator and regression estimator under our dependence structure.
Asymptotic results such as uniform convergence rate, $L^p$ convergence rate,
asymptotic normality are obtained under fairly mild conditions. In
particular, we allow certain long memory processes as well.
A closely related problem, the recursive kernel estimator where the
bandwidth changes with each observation, is still in progress and I hope to
get it done pretty soon. Therefore, I may talk about the recursive
estimator as well if time permits.
Feb. 9, 2011
Wei Sun :
3 p.m. in SEO 636
Abstract
K-means clustering is a widely used tool for cluster analysis due to its conceptual
simplicity and computational efficiency. However, its performance can be distorted
when clustering high-dimensional data where the number of variables becomes relatively large and many of them may contain no information about the clustering structure. In this talk we will discuss a novel high-dimensional cluster analysis method via regularized k-means clustering, which can simultaneously cluster similar observations and eliminate redundant variables. The key idea is to formulate the k-means clustering in a form of regularization, with an adaptive group lasso penalty term on cluster centers. Then we will talk about the selection criterion based on clustering stability to optimally balance the trade-off between the clustering model fitting and sparsity. The effectiveness of the proposed method is demonstrated through a variety of numerical experiments as well as applications to two gene microarray examples.
Feb. 16, 2011
Abhyuday Mandal :
3 p.m. in SEO 636
Abstract
We consider the problem of obtaining locally D-optimal designs
for factorial experiments with qualitative factors at two levels each with
binary response. For the 2^2 factorial experiment with main effects model
we obtain optimal designs analytically in special cases and demonstrate
how to obtain a solution in the general case using Cylindrical Algebraic
Decomposition. We also study the sensitivity of the D-optimal designs to
misspecification of the assumed parameter values. When there is no basis
to make an informed choice of the assumed values, we recommend the use of
the uniform design. For the general 2^k case we show that the uniform
design has a maximin property.
Feb. 23, 2011
Jie Yang :
3 p.m. in SEO 636
Abstract
We propose a stochastic classification model based on a permanent process. Unlike many research works in the literature, the proposed model assumes only exchangeability instead of independence on observations. Regardless of the number of classes or the dimension of the feature variables, the model may require only 2-3 parameters for fitting the covariance structure within clusters. It works well even if the class
occupies non-convex, disjoint regions, or regions overlapped with other
classes in the feature space. The proposed model requires calculation
of ratios of weighted permanents, which is an NP-hard problem. We propose
a series of approximations for weighted permanent ratio based on cyclic
expansions. The classification based on cyclic approximations works
reasonably well.
March 2, 2011
William Zhao :
3 p.m. in SEO 636
Abstract
An Overview of Pharmaceutical R&D and Roles of Statisticians. The drug research and
development process will be introduced. Examples will be provided. Experiences will
be shared on the roles of statisticians in the process.
March 16, 2011
Zhengyuan Zhu :
3 p.m. in SEO 636
Abstract
Laplace approximation has been used in many statistical applications
to approximate integrals. In this talk we present several applications
which use high order Laplace approximation to derive theoretical
results and develop efficient algorithms, including construction of
prediction interval which has zero second order coverage probability bias,
design criteria for prediction with estimated covariance parameters,
theoretical comparison of predictive densities for dependent observations,
and approximate inference for spatial generalized linear mixed models.
March 30, 2011
Robert Gibbons :
3 p.m. in SEO 636
Abstract
In 2003, the U.S. FDA, MHRA in the U.K., and European union released
public health advisories for a possible causal link between
antidepressant treatment and suicide in children and adolescents ages
18 and under. This led the U.S. FDA to issue a black box warning for
antidepressant treatment of childhood depression in 2004, which was
later extended to include young adults (18-24) in 2006. Following
these warnings, rather than observing the anticipated decrease in
youth suicide rates, record increases in youth suicide rates were
observed in both the U.S. and Europe. In this presentation, we
review the data and statistical methodology that led to the public
health advisories and black box warning, and the data that led to the
record increases in youth suicide rates and discuss their possible
relationship. New statistical and experimental design approaches to
post-marketing drug safety surveillance are developed, discussed and
illustrated.
April 6, 2011
Haiyan Cai :
3 p.m. in SEO 636
Abstract
In this talk I plan to give a brief introduction of the Division of
Mathematical Sciences (DMS) of NSF and its funding opportunities,
including its investment goals, investment areas, disciplinary and interdisciplinary
programs, and proposal evaluation criteria. I will also be happy to
answer questions.
April 13, 2011
Lulu Kang :
3 p.m. in SEO 636
Abstract
Computer experiments simulate the engineering systems by implementing the mathematical models governing the systems in computers. Recently, experiments having large number of input variables and experimental runs started to emerge. In the existing literature, kriging has been commonly used for approximating the complex computer models, but it has limitations for dealing with the large-scale experiments due to its computational complexity and numerical stability. In this work, I propose a new modeling approach known as regression-based inverse distance weighting (RIDW). The new predictor is shown to be computationally more efficient than kriging while producing comparable prediction performance. We also develop a heuristic method for constructing confidence intervals for prediction. I will also discuss extensions of RIDW and my future research directions on this exciting topic.
April 20, 2011
Min Yang :
3 p.m. in SEO 636
Abstract
In many practical studies, we need to consider the optimality problem
under variety constrains. For example the cost of experiment may depend on the
experiment point we choose and we need to control the total cost. On the other hand,
we may want to make sure the design is at least efficient at certain level under
different optimality criterion. In this talk, we will discuss such question.
April 27, 2011
Bikas K. Sinha :
3 p.m. in SEO 636
Abstract
In the context of a finite labeled population, for estimation
of a population total or mean, there does not exist any umvue, unless the
sampling design is trivially a cluster sampling design. This is the celebrated
Godambe's non-existence theorem. Basu established another version of it.
We revisit the proof and strengthen the result by using matrix arguments.
May 4, 2011
Reza Meshkani :
3 p.m. in SEO 636
Abstract
In this talk, the limitations of normal model for Analysis of Covariance for positive right-skewed variables are considered. Specifically, an Inverse Gaussian variable is considered whose variance depends on its mean thus violating the usual assumptions of Normal linear model. Instead of appealing to transformations which makes interpretations of the results awkward, we propose a method of direct statistical analysis from both Maximum Likelihood and Bayesian perspectives. The formulas for adjusting treatment effects are given and their properties are discussed. To provide explicit formulas, conjugate priors are considered. The posterior distributions are derived and procedures of adjustment for covariates are presented.
May 5, 2011
Ravindra Bapat :
3 p.m. in SEO 1227
Abstract
The max algebra consists of the set of real numbers,
along with negative infinity, equipped with two binary
operations, maximization and addition. The algebra is
useful in describing certain conventionally nonlinear
systems in a linear fashion. We discuss basic aspects
of the eigenproblem in max algebra, pointing out that
it can be seen as a limiting case of the Perron-Frobenius
theory for nonnegative matrices. We also discuss properties
of the permanent over the max algebra, which is the same as
the maximum value of the classical assignment problem.
Aug. 3, 2011
Prof Bikas K Sinha :
4 p.m. in SEO 636
Abstract
Considered is a multi-species assemblage in an infinite population
with unknown and possibly heterogeneous species' abundance levels. A random
sample of a fixed size (n) has been drawn only to realize a certain
species distribution. At this stage, there are two interesting inference
problems : (i) Prediction of the unknown proportion of the collective abundance
of all hitherto unrealized species; (ii) Assessment of the quantum of additional
units to be sampled [in terms of the sample size n] to realize a certain number
of hitherto unobserved species. Finite population analogue of this problem is also worth discussing.
I propose to review the available literature in this fascinating area of
research.
Aug. 24, 2011
Statistics Group :
4 p.m. in SEO 636
Aug. 31, 2011
Kashinath Chatterjee :
4 p.m. in SEO 636
Abstract
The main objective of this presentation is to introduce the notion of Search
Design, pioneered by Srivatava (1975), and its application to model selection
in fractional factorial experiments.
Sept. 7, 2011
Ryan Martin :
4 p.m. in SEO 636
Abstract
Testing if a p-dimensional sample, for p >= 1, comes from a normal population is a
fundamental problem in statistics. In this talk I will describe a new Bayesian test of p-variate
normality against an alternative hypothesis characterized by a certain Dirichlet process mixture
model. I will show that this nonparametric alternative satisfies the desirable embedding and
predictive matching properties with respect to the normal null model. To compute the Bayes
factor, an efficient sequential importance sampler is is proposed for evaluating the marginal
likelihood under the nonparametric alternative. Numerical examples demonstrate that the proposed
test has satisfactory discriminatory power when the distribution is not normal, and does not tend
to over-fit when the distribution is normal.
Sept. 14, 2011
Cheng Ouyang :
4 p.m. in SEO 636
Abstract
Stochastic differential equations (SDE) driven by various random
processes are important subject in both probability theory and applications, as
they provide mathematical models for systems that evolve under random forces.
Among them, study of SDE's driven by fractional Brownian motions is an active
area in current research. In the talk, I will first give a brief introduction
to this topic, and then present two resent results - namely, the concentration
property and Log-Sobolev inequality - on the law of solutions to SDE's driven
by fractional Brownian motions.
Sept. 21, 2011
Zhifan Zhang and Tu Xu :
4 p.m. in SEO 636
Abstract
(By Zhifan Zhang) Latent Class Analysis (LCA) is used to find groups or subtypes of cases in
multivariate categorical data. In Health care Market Research, Latent
Class Analysis is used in consumer segmentation to identify discriminating
variables that would enable the development of differentiated segments of
customers to inform post-Reform Strategies. In this talk, we will discuss
the usage of Latent Class Analysis and Regression analysis in Health care
Market Research.
(By Tu Xu) Multiclass probability estimation is an very important problem
in statistics and data mining. The traditional probability estimation
problem is commonly dealt by regression techniques such as multiple
logistic regression, or the density estimation approaches such as linear
discriminant analysis(LDA) and quadratic discriminant analysis. In this
talk,we propose a new model-free method for estimating multiclass
probabilities based quantile kernel regrssion will be introduced. It does
not impose any strong parametric assumption on the underlying distribution
and can be applied for a wide range of large-margin classification
methods. Our simulations show great performance of our method compared
with existing methods.
Sept. 28, 2011
Bo Li :
4 p.m. in SEO 636
Abstract
We propose a framework in light of the delay effect to model the
asymmetry of multivariate
covariance functions that is often exhibited in real data. This general
approach can
endow any valid symmetric multivariate covariance function with the
ability of modeling
asymmetry and is very easy to implement. Our simulations and real data
examples show
that asymmetric multivariate covariance functions based on our approach
can achieve
remarkable improvements in prediction over symmetric models.
Oct. 5, 2011
Mary Sara McPeek :
4 p.m. in SEO 636
Abstract
Common diseases such as asthma, diabetes, and hypertension,
which currently account for a large portion of the health
care burden, are complex in the sense that they are influenced
by many factors, both environmental and genetic. One fundamental
problem of interest is to understand what the genetic risk factors are
that predispose some people to get a particular complex disease.
Despite the potential for complex traits to have X-linked causal genes,
genetic association methods have primarily been developed for the analysis
of markers on the autosomal chromosomes, and significantly less attention
has been given to analyzing X-linked markers. We develop methods for
case-control association testing of X-chromosome variants in samples in
which some individuals are related. Our methods are applicable to
association studies with completely general combinations of family and
case-control designs, including large complex pedigrees. Even in the
context of large complex pedigrees, the methods are computationally
feasible for analysis of millions of variants. We allow for sex-specific
prevalence, and we allow both unaffected controls and controls of unknown
phenotype in the analysis. We discuss some of the distinct challenges
posed by X-chromosome association analysis in contrast to autosomal
association analysis. We discuss the performance of the methods
in the context of several data sets as well as in simulations.
Oct. 12, 2011
Ang Wei :
4 p.m. in SEO 636
Abstract
Unsigned Gaussian functions (functions involving absolute value)
arise in a variety of contexts, such as random polynomials and matrices. We
apply integral representations and matrix analysis to estimate the absolute
moments and quadratic forms of Gaussian vectors. We also present applications
of these results to game theory, communication theory and astrophysics.
Oct. 19, 2011
Tian Zheng :
4 p.m. in SEO 636
Abstract
Aggregated Relational Data (ARD) are indirect network data collected
using survey questions of the form "how many X's do you know?" It is
most often used to estimate the size of populations that are difficult
to count directly and allows researchers to choose specific
subpopulations of interest without sampling or surveying members of
these subpopulations directly. What has been under-utilized is the
indirect information on social structure captured by ARD. In this
talk, I present a latent space model and Bayesian computation
framework for inference and estimation of social structures using ARD
from non-network samples in social networks, the variation of social
structures in subnetworks, and the relations between (hard-to-reach)
subpopulations.
Oct. 25, 2011
Fabrice Baudoin :
4 p.m. in SEO 636
Abstract
In this talk we will review several results on stochastic
differential equations driven by fractional Brownian motions that the speakers
obtained in a series of more or less recent works. We shall in particular focus
on the study of gradients bounds, Gaussian heat kernels bounds and small time
asymptotics for the operators naturally associated with such equations. The
presentation will be based on joint works with L. Coutin, M. Hairer, C.
Ouyang and S. Tindel.
Oct. 26, 2011
Ivan Mizera :
4 p.m. in SEO 636
Abstract
Directional quantile envelopes---essentially, depth contours---are a
possible way to condense the directional quantile information, the
information carried by the quantiles of projections. In typical
circumstances, they allow for relatively faithful and straightforward
retrieval of the directional quantiles, offering a straightforward
probabilistic interpretation in terms of the tangent mass at smooth
boundary points. They can be viewed as a natural, nonparametric
extension of ``multivariate quantiles'' yielded by fitted multivariate
normal distribution, and, as illustrated on data examples, their
construction can be adapted to elaborate frameworks---like estimation
of extreme quantiles, and directional quantile regression---that
require more sophisticated estimation methods than simply evaluating
quantiles for empirical distributions. Their estimates are affine
equivariant whenever the estimators of directional quantiles are
translation and scale equivariant; mathematically, they express the
dual aspect of directional quantiles.
Nov. 2, 2011
Surya Tapas Tokdar :
4 p.m. in SEO 636
Abstract
I'll introduce a semi-parametric Bayesian framework for a
simultaneous analysis of linear quantile regression models. A simultaneous
analysis is essential to attain the true potential of the quantile regression
framework, but is computationally challenging due to the associated
monotonicity constraint on the quantile curves. For a univariate covariate, we
present a simpler equivalent characterization of the monotonicity constraint
through an interpolation of two monotone curves. The resulting formulation
leads to a tractable likelihood function and is embedded within a Bayesian
framework where the two monotone curves are modeled via logistic
transformations of a smooth Gaussian process. A multivariate extension is
proposed by combining the full support univariate model with a linear
projection of the predictors. The resulting single-index model remains easy to
fit and provides substantial and measurable improvement over the first order
linear heteroscedastic model. I'll provide two illustrative applications to
tropical cyclone intensity and birth weight.
Nov. 9, 2011
Elton Hsu :
4 p.m. in SEO 636
Abstract
We will show how to use synchronizing Brownian motion to prove
Talagrand's transportation cost inequality for the standard Gaussian measure. This proof can be
readily generalized to a generalization of this inequality to the heat kernel measure on a
Riemannian manifold using the synchronizing coupling of Riemannian Brownian motion.
Nov. 16, 2011
Hanxiang Peng :
4 p.m. in SEO 636
Abstract
In this talk, I will focus on maximum empirical likelihood
estimation in the case of constraint functions that may be
discontinuous and/or depend on additional parameters.
The later is the case in applications to semiparametric models
where the constraint functions may depend on the nuisance parameter.
Our results are thus formulated for empirical likelihoods based on
estimated constraint functions that may also be irregular.
Applications of our results are discussed to inference problems
about quantiles under possibly additional information on the
underlying distribution and to partial adaption.
Nov. 30, 2011
Surajit Borkotokey :
4 p.m. in SEO 636
Abstract
We propose an allocation rule that takes into account the importance of
players and their links.
Since a network describes the interaction structure between agents, our
allocation rule covers both bilateral and multilateral interactions.
We provide a characterization of this rule in terms of well known axioms
and compare it to other allocation rules in the literature.
Jan. 11, 2012
Shmuel Friedland :
4:15 p.m. in SEO 636
Abstract
In many instances in measuring multidimensional data, as matrices and tensors, one confronts the following
problems: noisy data, missing entries and data reduction.
There are many statistical and mathematical methods to deal with these problems.
In this talk we survey some of the known methods and expand on the methods that the speaker was working on.
A variant of this talk is available at
http://homepages.math.uic.edu/$\sim$friedlan/complmattenSep11.pdf
Jan. 25, 2012
Yan Sun, Ella Revzin :
4 p.m. in SEO 636
Abstract
Yan Sun's Abstract:
Soybean is an important crop in the US and around the world. So the analysis of soybean yield components as well as the growth components is always a hot topic in research. Our work originates from the requirement of a client and is based on the experiment data provided by him. The experiment has a split-split plot structure. Our task is to find significant factors for seed yield and protein concentration, which are both yield components. Also, we establish regression functions for some growth components (leaf area and leaf dry biomass) with respect to the growing time.
Ella Revzin's Abstract:
In the Fall 2011, the University of Illinois at Chicago changed its
mathematics placement policy. Instead of placing students based on
ACT scores and results on written tests, the mathematics department
switched to using an online testing and learning software, ALEKS.
This talk will summarize the results from analyzing student outcomes
after and before the policy change; as well as an assessment of how
effective the ALEKS system was in placing students in selected
undergraduate courses.
Feb. 1, 2012
Jin Zhang :
4:15 p.m. in SEO 636
Abstract
We estimate the daily integrated variance and covariance of stock returns using high-frequency data in the presence of jumps, market microstructure noise and non-synchronous trading. For this we propose jump robust two time scale (co)variance estimators and verify their reduced bias and mean square error in simulation studies. We use these estimators to construct the ex-post portfolio realized volatility (RV) budget, determining each portfolio component's contribution to the RV of the portfolio return. These RV budgets provide insight into the risk concentration of a portfolio. Furthermore, the RV budgets can be directly used in a portfolio strategy, called the equal-risk-contribution allocation strategy. This yields both a higher average return and lower standard deviation out-of-sample than the equal-weight portfolio for the stocks in the Dow Jones Industrial Average over the period October 2007-May 2009.
Feb. 8, 2012
Yue Yu :
4 p.m. in SEO 636
Abstract
Study of measuring agreement is mainly aimed to answer one
question, whether the readings from one instrument/method agree with
the ones from another instrument/method. In this talk, we are going to
present a general method to assess agreement for a wide range of data
types with repeated measurements using linear and generalized linear
mixed models. Likelihood-based approaches are developed to estimate
all the within- and between-instrument agreement statistics. and
asymptotic properties of these agreement estimates are discussed for
different data structures. Furthermore, our method has the merit of
handling missing values and covariates naturally. And a new set of
restricted agreement statistics is proposed in order to capture the
true random variations and between-instrument effects rather than the
covariate effects. Simulations and several case studies, involving
method comparison and bioequivalence, are used to show the accuracy
and effectiveness of our method.
Feb. 15, 2012
Chenglong Yu :
4 p.m. in SEO 636
Abstract
Among all existing alignment-free methods for comparing biological
sequences, the sequence graphical representation provides a simple approach to
view, sort, and compare gene structures. The aim of graphical representation is
to display DNA or protein sequences graphically so that we can easily find out
visually how similar or how different they are. Of course, only the visual
comparison of sequences is not enough for the follow-up research work. We need
more accurate comparison. This leads us to develop the application of the
graphical representation for biological sequences. I will talk about two
contributions for this direction. (1) We construct a protein map with the help
of our proposed new graphical representation for protein sequences. Each
protein sequence can be represented as a point in this map, and cluster
analysis of proteins can be performed for comparison between the points. This
protein map can be used to mathematically specify the similarity of two
proteins and predict properties of an unknown protein based on its amino acid
sequence. (2) We construct a novel genome space with biological geometry, which
is a subspace in R^N. In this space each point corresponds to a genome. The
natural distance between two points in the genome space reflects the biological
distance between these two genomes. The genome space will provide a new
powerful tool for analyzing the classification of genomes and their
phylogenetic relationships.
Feb. 21, 2012
Bikas Sinha :
4 p.m. in SEO 636
Abstract
de la Garza Phenomenon relates to the Information Matrix in
the context of a standard Gauss-Markov Linear Model. It
works well in the framework of approximate or continuous designs.
For discrete designs, one has to be careful in extracting
its full spirit. We propose to discuss some features of this highly
fascinating area of research.
Feb. 22, 2012
Yixin Fang :
4 p.m. in SEO 636
Abstract
In family studies with multiple continuous phenotypes, we are interested
in finding linear combinations of the phenotypes with large
heritabilities, which can be considered as new phenotypes for genetic
analysis. The problem can be recast as linear discriminant analysis (LDA).
When the number of phenotypes is large, LDA is not appropriate for two
reasons: the standard estimate for the within-family covariance matrix is
singular, and it is difficult to interpret the newly defined phenotypes.
Here we propose a novel version of penalized LDA, with an $L_1^2$ penalty
in the denominator of the Rayleigh quotient. Besides overcoming the above
two problems, the proposed method has at least three advantages compared
with the existing regularization methods. First, it solves the singularity
problem and achieves the sparsity property simultaneously. Second, the
method is scale-invariant. Third, the consistency can be proved. We
evaluate the performances of the method using simulations and two family
studies.
Feb. 29, 2012
Jennifer Pajda-Delao / Julien Leider :
4 p.m. in SEO 636
Abstract
Jennifer's Abstract:
We evaluate the status and distribution of swamp rabbits (Sylvilagus
aquaticus) in Missouri using the Incidence Function Model and logistic
regression in an effort to assess the long term viability of the
Missouri metapopulation. We used results of latrine surveys performed
in 1992 and 2001 to estimate the likelihood of persistence of swamp
rabbits over periods of 9 to 1000 years. Under current conditions, more
than 50% of the patches are predicted to contain rabbits after 1000
years. Logistic regression revealed that both patch area and patch
isolation were significantly related to patch occupancy, and play key
roles in the incidence of swamp rabbits.
Julien's Abstract:
This study uses quantile regression combined with time series methods to
analyze change in temperatures in Chicago during the period 1960-2010.
It builds on previous work in applying quantile regression methods to
climate data by Timofeev and Sterin (2010) and work by the Chicago Climate
Task Force on analyzing climate change in Chicago. We use data from the
Chicago O'Hare Airport weather station archived by the National Climatic
Data Center to look at changes in weekly average temperatures. We use the
method described by Xiao et al. (2003) to remove autocorrelation in the
data, the rank-score method with IID assumption to calculate confidence
intervals, and nonparametric local linear quantile regression to estimate
temperature trends. We find that the decade 1960-1969 was significantly
cooler than later decades around the middle of the yearly seasonal
cycle at both the median and 95th percentile of the temperature
distribution. However, we do not find a significant change across later
decades of the study period.
March 7, 2012
Chao Zhu :
4 p.m. in SEO 636
Abstract
We consider the optimal harvesting strategy for a single species living in
random environments whose growth is given by a regime-switching diffusion. Harvesting
acts as a (stochastic) control on the size of the population. The objective is to
find a harvesting strategy which maximizes the expected total discounted income from
harvesting up to the time of extinction of the species; the income rate is allowed to
be state- and environment-dependent. This is a singular stochastic control problem
with both the extinction time and the optimal harvesting policy depending on the
initial condition. One aspect of receiving payments up to the random time of
extinction is that small changes in the initial population size may significantly
alter the extinction time when using the same harvesting policy. Consequently, one no
longer obtains continuity of the value function using standard arguments for either
regular or singular control problems having a fixed time horizon. We introduce a new
sufficient condition under which the continuity of the value function for the
regime-switching model is established. Further, it is shown that the value function
is the unique viscosity solution of a coupled system of quasi-variational
inequalities. We also establishes a verification theorem and, based on this theorem,
an $\varepsilon$-optimal harvesting strategy is constructed under certain conditions
on the model. This is a joint work with Qingshuo Song and Richard Stockbridge.
March 14, 2012
Juan Du :
4 p.m. in SEO 636
Abstract
We derive several classes of covariance matrix functions whose entries are
compactly supported. These compactly supported matrix functions are used as building
blocks to formulate other covariance matrix functions for modeling of multivariate
spatial processes. In particular, a multivariate version of the celebrated spherical
model is produced, as well as a class of second-order multivariate stochastic processes
whose direct and cross covariance functions are of Pólya type. On the other hand, by
employing some of the proposed compactly supported correlation matrix functions as the
tapering matrix function, we study the multivariate generalization of the spatial
covariance tapering technique, which is useful to mitigate the numerical burdens in
dealing with the large spatial data sets by making covariance matrices sparse.
Simulation study is conducted to show the computational efficiency and application in
spatial prediction by using proposed multivariate tapering technique.
March 28, 2012
Liming Feng :
4 p.m. in SEO 636
Abstract
Analytic characteristic functions naturally arise in financial
engineering applications. We explore the analyticity of such characteristic
functions and propose simple but highly accurate inversion schemes. The schemes
have the following advantages: (1) they are very easy to implement; one does not
need to rely on commercial numerical packages; (2) despite the simplicity, they
are highly accurate, with exponentially decaying errors; (3) they admit explicit
error estimates that only depend on the given characteristic function; (4)
multiple values of the desired quantity can be computed simultaneously using the
fast Fourier transform. We illustrate the effectiveness of the schemes with
financial engineering examples, including the valuation of options with barrier,
lookback and early exercise features in Lévy models, as well as Monte Carlo
simulation of Lévy processes with analytic characteristic functions.
April 4, 2012
Bhaskara Rao Kopparty :
4 p.m. in SEO 636
Abstract
Yes. That is true. All of us know that some real and complex matrices have inverses. Not all of them. But we also know that all real and complex matrices have generalized inverses. These are extensively used by statisticians. What about matrices that have only integer entries? Would such a matrix have a generalized inverse whose entries are all integers? What if a matrix has all entries as polynomials? We shall discuss various questions about generalized inverses and see some exciting results.
Here is one of the results: An integer matrix of rank r has a generalized inverse if and only if the greatest common divisor of all r x r minors of A is 1.
April 11, 2012
Steven P. Lalley :
4 p.m. in SEO 636
Abstract
I will survey of some recent work in scaling limits of stochastic spatial SIR epidemics. In these models,
colonies of $N$ individuals are located at lattice points of $\mathbb{Z}^d$. Each individual is, at any time, <i>susceptible</i>, <i>infected</i>,
or <i>recovered</i> (and immune to future infection). Infected individuals recover at rate 1, and infect susceptibles
in the same or neighboring colonies at rate (say) $\lambda/N$. When $\lambda = 1/(2d + 1)$, where $d$ is the dimension of the
lattice, the epidemic is <i>critical</i>: the mean number of new infections produced by a single infected individual
when all other individuals in the same or neighboring colonies are susceptible is 1. The questions of natural
interest center on the duration and spatial extent of a critical epidemic initiated by a large number $N^\alpha$
(where $0<\alpha <1$) of infected individuals all located at the colony at the origin of the lattice. The main results are
large$-N$ <i>scaling laws</i> for the epidemic.
The stochastic epidemic models can be reformulated as <i>percolation processes</i> on graphs that are, in a natural
sense, hybrids of the standard regular lattices and the complete graph on $N$ vertices. The spread of the epidemic
is determined by the geometry of the connected clusters of the associated percolation process. The scaling laws
for epidemic processes consequently translate to scaling laws for percolation clusters.
April 18, 2012
Liang Hong :
4 p.m. in SEO 636
Abstract
Contingent Means are ubiquitous in modern finance and actuarial science. However,
Standard textbooks on actuarial science or statistics do not elaborate on the correct
interpretation of contingent means, leaving the actuaries at risk of making a blunder.
In this talk, we will give the correct interpretation of contingent means both
heuristically and theoretically so that one will be aware of some common
misconceptions and avoid pitfalls in their work. We will also discuss the applications
of contingent means in insurance and quantitative finance.
April 25, 2012
Jin Feng :
4 p.m. in SEO 636
Abstract
The vorticity formulation of 2-D incompressible Navier-Stokes equation can be
viewed
as mean-field limit of stochastic interacting point vortices. As number of
particles goes
to infinity and viscosity term goes to zero, we arrive at inviscid limit of 2-D
incompressible
Euler equation.
We study multi-scale large deviation limits of such model, on the torus, as
particle number
and time go large but viscosity goes small. The result gives a first principle
approach to
establish the Onsager-Joyce-Montgomery theory as limit theorem derived from
microscopically
defined non-equilibrium models. The Onsager-Joyce-Montgomery theory concerns large
time coherent structures of vortex dynamics associated with 2-D Euler equation. It
was previously
informally formulated using equilibrium models only.
The talk is based on joint works (some of which are ongoing) of the speaker with
Fasto Gozzi,
Tom Kurtz and Andrzej Swiech, a SQuaRE team funded by American Institute of
Mathematics.
June 27, 2012
Qingshuo Song :
3 p.m. in SEO 636
Abstract
We study the portfolio problem of maximizing the out-performance probability over a random
benchmark through dynamic trading with a fixed initial capital. Under a general incomplete
market framework, this stochastic control problem can be formulated as a composite pure
hypothesis testing problem. We analyze the connection between this pure testing problem and
its randomized counterpart, and from latter we derive a dual representation for the maximal
outperformance probability. Moreover, in a complete market setting, we provide a closed-form
solution to the problem of beating a leveraged exchange traded fund. For a general benchmark
under an incomplete stochastic factor model, we provide the Hamilton-Jacobi-Bellman PDE
characterization for the maximal out-performance probability. It's a joint work with
Tim Leung and Jie Yang.
Aug. 29, 2012
Statistics Group :
4 p.m. in SEO 636
Sept. 5, 2012
Prof. Ryan Martin :
4 p.m. in SEO 636
Abstract
In the frequentist program, inferential methods with exact control on error rates are a primary focus. Methods based on asymptotic distribution theory may not be suitable in a particular problem, in which case, a numerical method is needed. In this talk I shall present a general, yet simple, Monte Carlo-driven framework for the construction of frequentist procedures based on plausibility functions. It is proved that the suitably defined plausibility function-based tests and confidence regions have desired frequentist properties. Moreover, in an important special case involving likelihood ratios, conditions are given such that the plausibility function behaves asymptotically like a consistent Bayesian posterior distribution. An extension of the proposed method is also given for the case where nuisance parameters are present. I shall give several examples to illustrate the method's flexibility and to demonstrate its performance compared to existing numerical and analytical methods.
Sept. 12, 2012
Matthew Bourque (PhD Candidate) :
4 p.m. in SEO 636
Abstract
Stochastic games model a competitive situation between two players in discrete time steps over an infinite horizon, in which players' payoffs at each stage depend on both players' action choice. They can be seen as generalizations of both repeated games and Markov decision processes (MDPs). Policy improvement algorithms are an important category of fast algorithms for solving MDPs. In this talk, we will give an introduction to stochastic games, in particular zero-sum games of perfect information and with ARAT structure, and discuss a policy improvement algorithm for finding optimal policies for both players for these categories of games when players are evaluating their payoff steams via a limiting average.
Sept. 19, 2012
Dr. Stephen G. Eick :
4 p.m. in SEO 636
Abstract
The next decade will be an ideal time to become a Statistical Entrepreneur. We are currently at the beginning edge of what will be a massive deployment of sensors, GPS-enabled smartphones, ubiquitous video, and all sorts of smart devices. This trend, often called "The Internet of Things," will involve devices which stream wireless real-time data back into cloud-based servers that will perform analytical calculations. For statisticians, we are entering a data rich environment while there will be tremendous need to create tools to analyze and make sense of this data.
After spending nearly 15 years working in academia and for big business, I became a Statistical Entrepreneur. I have been involved with 1⁄2 a dozen emerging growth analytics software companies would like to share some of my experiences. This is by far the best time that I have ever seen for statistical entrepreneurship.
Sept. 26, 2012
Prof. Anindya Bhadra :
4 p.m. in SEO 636
Abstract
We describe a Bayesian technique to (a) perform a sparse joint selection of significant predictor variables and significant inverse covariance matrix elements of the response variables in a high-dimensional linear Gaussian sparse seemingly unrelated regression (SSUR) setting and (b) perform an association analysis between the high-dimensional sets of predictors and responses in such a setting. To search the high-dimensional model space, where both the number of predictors and the number of possibly correlated responses can be larger than the sample size, we demonstrate that a marginalization-based collapsed Gibbs sampler, in combination with spike and slab type of priors, offers a computationally feasible and efficient solution. As an example, we apply our method to an expression quantitative trait loci (eQTL) analysis on publicly available single nucleotide polymorphism (SNP) and gene expression data for humans where the primary interest lies in finding the significant associations between the sets of SNPs and possibly correlated genetic transcripts. Our method also allows for inference on the sparse interaction network of the transcripts (response variables) after accounting for the effect of the SNPs (predictor variables). We exploit properties of Gaussian graphical models to make statements concerning conditional independence of the responses. Our method compares favorably to existing Bayesian approaches developed for this purpose. This is joint work with Bani K. Mallick of Texas A&M University.
Oct. 3, 2012
Dr. XiaoHui Chen :
4 p.m. in SEO 636
Abstract
Covariance matrix and its inverse (a.k.a. precision matrix) play a central role in a
broad range of problems in statistics and machine learning. In the past few years, there
has been an explosion of interest in regularized covariance and precision matrix
estimation for high-dimensional i.i.d. random vectors with sub-Gaussian tails.
In this talk, we shall discuss the estimation of covariance and precision matrices
for stationary and locally stationary high-dimensional time series. In the latter case,
the covariance matrices evolve smoothly in time and thus form a covariance matrix
function. Under the framework of Wu (2005)'s functional dependence measure, we obtain
the rate of convergence for the thresholded covariance matrix estimate and illustrate
how the dependence affects the rate of convergence. Asymptotic properties are also
obtained for the precision matrix estimate based on the graphical Lasso principle.
Our theory substantially generalizes earlier ones by allowing dependence, by allowing
non-stationarity and by relaxing the associated moment conditions. Our new results
have implications on a number of classical problems, including spatial-temporal
statistics and graphical models, among many others.
Oct. 10, 2012
Prof. Dennis Lin :
4 p.m. in SEO 636
Abstract
In the past decades, we have witnessed the revolution of information technology. Its impact to statistical research is enormous. This talk attempts to address some recent developments and potential research issues in Business, Industry and Government (BIG) Statistics, with special focus on computer experiment and information systems.
An overall introduction and review will be given, followed by specific research potentials. Some initial results will be presented, and future research problems will be suggested. If time permits, I will also discuss some recent advances in Search Engine and RFID study. Slides of this talk can be downloaded at the website
http://www.personal.psu.edu/users/j/x/jxz203/lin/Lin_pub/
Oct. 24, 2012
Prof. Shuva Gupta :
4 p.m. in SEO 636
Abstract
Here we investigate two problems concerning the asymptotic properties of an $l_{1}$ penalized regression. In the first problem we study the asymptotic distribution of the Lasso estimator for regression models with dependent errors. The asymptotic distribution of the Lasso estimator for regression models with independent errors has been investigated by Knight and Fu. Here we extend these results to regression models with a general weak dependence structure. We determine the asymptotic distribution of the Lasso estimator when the number of parameters M is fixed and the number of observations, n, converges to $\infty$. We show that, for an appropriate choice of the tuning parameter of the method, this asymptotic distribution reduces to a multivariate normal distribution. We also provide some illustrative examples. In the second problem, we deal with the asymptotic distribution of residual empirical process of residuals in an adaptive lasso setting. We study the asymptotic properties of the empirical residual process and then use it to investigate the problem of goodness of fit when p<n but increases with n. We explore different applications of residual empirical process eg test of goodness of fit. This work was largely motivated by that of Chen and Lockhart (2001).
Oct. 31, 2012
Dr. Ed Vonesh :
4 p.m. in SEO 636
Abstract
Correlated response data, either discrete (nominal, ordinal, counts), continuous or a combination thereof, occur in numerous disciplines and more often than not, require the use of statistical models that are nonlinear in the parameters of interest. Such models include generalized linear and generalized nonlinear models both of which can be further classified according to whether they are marginal or mixed-effects models. In this talk we briefly describe the different types of correlated response data and models encountered in practice. As some of the models can be quite complicated, there will often appear to be certain modeling limitations with available software. For users of SAS, such limitations would appear to include 1) how to conduct likelihood-based inference for nonlinear mixed-effects models with intra-subject correlation; 2) how to fit nonlinear mixed-effects models to data assuming non-Gaussian random effects; and 3) how to fit marginal generalized linear models to correlated response data using second order generalized estimating equations or maximum likelihood estimation. The focus of this talk will be on illustrating how one can fit mixed-effects models in SAS when the random effects are non-Gaussian. Following the work of Nelson et. al. (2006), the approach entails applying probability integral transformations when evaluating an integrated log-likelihood function. This technique is illustrated through an application that requires one to jointly model two dependent variables using a shared non-Gaussian random effect. It requires the user to be familiar with nonlinear mixed-effects models in general and also with how one can fit such models using the SAS procedure NLMIXED.
Nov. 7, 2012
Prof. John Stufken :
4 p.m. in SEO 636
Abstract
Identifying optimal designs for nonlinear models is a challenging problem for a variety of reasons. While the problem has received considerable attention over the last decades, a new approach has facilitated significant progress during the last three years in the context of local optimality. Although locally optimal designs are not always the preferred choice for applications, they are useful as a benchmark for other designs. In this talk we will cover some of the challenges related to identifying optimal designs for nonlinear models, discuss the new approach for identifying locally optimal designs that is based on finding small complete classes of designs, and illustrate the power of the approach through examples.
Nov. 14, 2012
Prof. Xuming He :
4 p.m. in SEO 636
Abstract
Quantile regression is semiparametric in the sense that no parametric likelihood is assumed in the model. A working likelihood can be used, but the resulting posterior may not have the validity for statistical inference. In this talk we will introduce Bayesian empirical likelihood for quantile regression, and show that it leads to asymptotically valid posterior inference. In addition, this approach enables us to make use of commonality across quantiles to improve efficiency of quantile estimation in data sparse areas. We will also introduce a notion of shrinking priors, and demonstrate how this new framework can help explain the efficiency gains of the Bayesian empirical likelihood method over the usual quantile estimates. The talk is based on joint work with Yunwen Yang (Drexel University).
Nov. 21, 2012
Prof Yimin Xiao :
4 p.m. in SEO 636
Abstract
By applying methods for studying Gaussian random fields, we investigate various analytic and fractal properties of the solutions of the stochastic heat equations driven by space-time white noise or fractional colored noise. These include fractal dimensions, exact modulus of continuity, hitting probabilities and existence of intersections. The proofs of these results are based on the properties of strong local nondeterminism. (Based on the on-going joint works with R. Dalang, C. Mueller and C. Tudor.)
Nov. 28, 2012
Prof. Malgorzata Bogdan :
4 p.m. in SEO 636
Abstract
In Bogdan et al (Ann. Statist. 2011) the asymptotic framework for the analysis of the Bayes risk of the multiple testing procedures under sparsity is proposed. Within this framework the rule is called Asymptotically Bayes Optimal under Sparsity (ABOS) if the ratio of its risk and the risk of the Bayes oracle converges to 1 as the number of tests, $m$, diverges to infinity and the proportion of alternatives among all tests, $p$, converges to zero. In Bogdan et al (2011) and Neuvial and Roquain (Ann. Statist., to appear) the conditions under which the popular Benjamini-Hochberg and Bonferroni procedures are ABOS are provided. We will discuss these results and provide an extension to the situation where the sample size $n$ used to calculate each of the test statistics goes to infinity with the number of tests $m$. We show that under mild restrictions on the loss function and the distribution of the magnitude of true signals a nontrivial asymptotic inference is possible only if $n$ increases to infinity at least at the rate of $\log m$. Based on this assumption precise conditions are given under which the Bonferroni correction with nominal Family Wise Error Rate (FWER) level $\alpha$ and the Benjamini-Hochberg procedure (BH) at FDR level $\alpha$ are asymptotically optimal.
In the second part of this talk these optimality results are carried over to model selection in the context of multiple regression with orthogonal regressors. Several modifications of Bayesian Information Criterion are considered, controlling either FWER or FDR, and conditions
are provided under which these selection criteria are ABOS. Finally the performance of the multiple testing rules and the model selection criteria is examined in a brief simulation study.
Dec. 5, 2012
Prof. Runze Li :
4 p.m. in SEO 636
Abstract
Ultra-high dimensional data often display heterogeneity due to either heteroscedastic variance or other forms of non-location-scale covariate effects. To accommodate heterogeneity, we advocate a more general interpretation of sparsity which assumes that only a small number of covariates influence the conditional distribution of the response variable given all candidate covariates; however, the sets of relevant covariates
may differ when we consider different segments of the conditional distribution. In this talk, I first introduce recent development on the methodology and theory of nonconvex penalized quantile linear regression in ultra-high dimension. I further propose a two-stage feature screening and cleaning procedure to study the estimation of the index parameter in heteroscedastic single-index models with ultrahigh dimensional covariates.
Sampling properties of the proposed procedures are studied. Finite sample performance of the proposed procedure is examined by Monte Carlo simulation studies. A real example example is used to illustrate the proposed methodology.
Jan. 23, 2013
Prof. SIAVASH H. SOHRAB :
4:15 p.m. in SEO 636
Abstract
A scale invariant model of statistical mechanics is applied to derive invariant forms of
conservation equations. A modified form of Cauchy stress tensor for fluid is presented that leads to modified Stokes assumption thus a finite coefficient of bulk viscosity. The phenomenon of Brownian motion is described as the state of equilibrium between suspended particles and molecular clusters that themselves possess Brownian motion. Physical space or Casimir vacuum is identified as a tachyonic fluid that is "stochastic ether" of Dirac or "hidden thermostat" of de Broglie, and is compressible in accordance with Planck's compressible ether. The stochastic definitions of Planck h and Boltzmann k constants are shown to respectively relate to the spatial and the temporal aspects of vacuum fluctuations. Hence, a modified definition of thermodynamic temperature is introduced that leads to predicted velocity of sound in agreement with observations. Also, a modified value of Joule-Mayer mechanical equivalent of heat is identified as the universal gas constant and is called De Pretto number 8338 which occurred in his mass-energy equivalence equation. Applying Boltzmann's combinatoric methods, invariant forms of Boltzmann, Planck, and Maxwell-Boltzmann distribution functions for equilibrium statistical fields including that of isotropic stationary turbulence are derived. The latter is shown to lead to the definitions of (electron, photon, neutrino) as the mostprobable equilibrium sizes of (photon, neutrino, tachyon) clusters, respectively. The physical basis for the coincidence of normalized spacings between zeros of Riemann zeta function and the normalized Maxwell-Boltzmann distribution and its connections to Riemann Hypothesis are examined. The zeros of Riemann zeta function are related to the zeros of particle velocities or "stationary states" through Euler's golden key thus providing a physical explanation for the location of the critical line. It is argued that because the energy spectrum of Casimir vacuum will be governed by Schrödinger equation of quantum mechanics, in view of Heisenberg matrix mechanics physical space should be described by noncommutative spectral geometry of Connes. Invariant forms of transport coefficients suggesting finite values of gravitational viscosity as well as hierarchies of vacua and absolute zero temperatures are described. Some of the implications of the results to the problem of thermodynamic irreversibility and Poincaré recurrence theorem are addressed. Invariant modified form of the first law of thermodynamics is derived and a modified definition of entropy is introduced that closes the gap between radiation and gas theory. Finally, new paradigms for hydrodynamic foundations of both Schrödinger as well as Dirac wave equations and transitions between Bohr stationary states in quantum mechanics are discussed.
Feb. 6, 2013
Prof. Ryan Martin :
4 p.m. in SEO 636
Abstract
In high-dimensional problems, the parameter of interest is called sparse if most of its components are zero. An important example is in high-dimensional regression, were it is believed that only a few of the many predictor variables explain variation in the response. Since we don't know beforehand which coordinates of the parameter vector are zero, the challenge is to simultaneously identify those which are zero and accurately estimate those which are non-zero. From a Bayesian point of view, to accommodate sparsity, an intuitive strategy is to consider a discrete-continuous mixture prior which allows coordinates to be exactly zero with positive probability. The relevant asymptotic theory looks for conditions such that the posterior distribution concentrates around the true signal at the best rate. In the first part of the talk, I will discuss some results along these lines from the very-recent literature. One drawback to the discrete-continuous mixture priors is that computation can be very difficult, so it is natural to ask if similar posterior concentration results can be achieve with computationally simpler non-mixture priors. This question is almost completely open, and the second part of the talk will discuss some aspects of this problem and what I think can be done.
Feb. 13, 2013
Prof Song Renming :
4 p.m. in SEO 636
Abstract
A subordinate Brownian motion can be obtained by replacing the time
parameter of a Brownian motion by an increasing Levy process (i.e.,
subordinator). Subordinate Brownian motions are very important in
various applications. In this talk, I will give a survey of some recent
results in the study of subordinate Brownian motions. In particular, I
will present results on sharp two-sided estimates of the Dirichlet
heat kernel estimates of subordinate Brownian motions.
Feb. 20, 2013
Gerald (Jerry) Phillips :
4 p.m. in SEO 636
Abstract
The non-clinical field in the healthcare industry is an important area that primarily consists of R&D, manufacturing and product field maintenance. Statistical applications in these areas will be demonstrated. In addition, skills required to provide successful statistical consulting will be discussed.
Feb. 27, 2013
Prof. Marlos Viana :
4 p.m. in SEO 636
Abstract
The seminar will present an overview of Fourier analysis over the dihedral groups as a method of determining (dihedral) orbit invariants as statistical summaries of (dihedral) experiments. While classical commutative harmonic analysis has a long history in science in general and in optics and vision studies in particular, the dihedral groups are certainly among the finite non-commutative groups of broadest application in those areas. The data-analytic aspects to be discussed include the notions of dihedral experiments, labeling arbitrariness of group orbits, determination and interpretation of the Fourier transforms as orbit invariants and statistical summaries. Some of the applications to briefly illustrated include the modeling of corneal curvature and power surfaces, polarimetric-enhanced retinal imaging methods, dihedral polynomial methods for wave-front aberration analysis, visual field decompositions, and visual perception studies. The theory and methods of dihedral analysis are also applicable in the studies of symbolic sequences in structural/functional molecular biology, vibrational spectroscopy, and several other fields, thus leading to a rich interplay of several basic disciplines such as physics, biology, algebra and statistics among others.
March 6, 2013
Prof. Jian Zou :
4 p.m. in SEO 636
Abstract
Portfolio allocation is one of the most fundamental problems in finance. The process of determining the optimal mix of assets to hold in the portfolio is a very important issue in risk management. It involves dividing an investment portfolio among different assets based on the volatilities of the asset returns. In the recent decades, it gains popularity to estimate volatilities of asset returns based on high-frequency data in financial economics. However the most available methods are not directly applicable when the number of assets involved is large, since small component-wise estimation errors could accumulate to large matrix-wise errors. This paper starts with a review on portfolio allocation and high-frequency financial time series. Then we introduce a new methodology to carry out efficient asset allocations using regularization on estimated integrated volatility via intra-day high-frequency data. We illustrate the methodology with the high-frequency price data on stocks traded in New York Stock Exchange over a period of 209 days in 2010. The theory and numerical results show that our approach perform well in portfolio allocation by pooling together the strengths of regularization and estimation from a high-frequency finance perspective.
March 13, 2013
Prof. Fabrice Baudoin :
4 p.m. in SEO 636
Abstract
This talk investigates several properties related to densities of solutions $(X_t)_{t\in[0,1]}$ to differential equations driven by a fractional Brownian motion with Hurst parameter $H>1/4$. We first determine conditions for strict positivity of the density of $X_t$. Then we obtain some exponential bounds for this density when the diffusion coefficient satisfies an elliptic type condition. Finally, still in the elliptic case, we derive some bounds on the hitting probabilities of sets by fractional differential systems in terms of Newtonian capacities.
March 20, 2013
Dr. Jennifer Hill :
4 p.m. in SEO 636
April 3, 2013
Prof. Yehua Li :
4 p.m. in SEO 636
Abstract
In disease surveillance applications, the disease events are modeled by spatial-temporal point processes. We propose a new class of semiparametric generalized linear mixed Cox model for such data, where the event rate is related to some known risk factors and some unknown latent random effects. We model the latent spatial-temporal process as spatially correlated functional data, and propose composite likelihood methods based on spline approximation to estimate the mean and covariance of the latent process. By performing functional principal component analysis to the latent process, we gain deeper understanding of the correlation structure in the point process, and we propose an empirical Bayes method to predict the latent spatial random effects, which can help highlighting the high risk spatial regions for the disease. Under an increasing domain and increasing knots asymptotic framework, we provide the asymptotic distribution for the parametric components in the model and the asymptotic convergence rate for the functional principal component estimators. We illustrate the methodology through a simulation study and an application to the Connecticut Tumor Registry data.
April 10, 2013
Prof. Peter Qian :
4 p.m. in SEO 636
Abstract
Gaussian process (GP) models are widely used in statistics, optimization, machine learning and other fields. Fitting a GP model with massive data is not only a challenge but also a mystery. On one hand, the nominal accuracy of a GP model is supposed to increase with the number of data points. On the other hand, fitting such a model to a large number of points encounters numerical singularity. To reconcile this contradiction, I will present a method to achieve both numerical stability and theoretical accuracy in fitting a massive GP model. This method obtains nested subsamples of the data, builds submodels for different subsets and then combines these models together to form an accurate prediction model. A decomposition of the overall model error into nominal and numeric portions is introduced to shed light on the theoretical underpinnings of the method. Bounds on the numeric and nominal error are developed to show that substantial gains in overall accuracy can be attained with this sequential method. Efficient algorithms are introduced to generate the required nested subsamples of the developed method.
April 17, 2013
Prof. Stephen Walker :
4 p.m. in SEO 636
Abstract
There are desirable models with attractive features which suffer from the problem of intractable normalizing constants. A general strategy for dealing with such constants in a Bayesian setting is presented. A range of illustrations, from the Fisher-Bingham distribution to power likelihood models, regression and time series models will be provided.
April 24, 2013
Prof. Zhen-Qing Chen :
4 p.m. in SEO 636
Abstract
Image an insect moves randomly in a plane with an infinite pole installed on it. In this talk, we will introduce and discuss Brownian motion on a state space with varying dimension. We will derive sharp two-sided estimates on its transition density function (also called heat kernel). The two-sided estimates is of Guassian type but the parabolic Harnack inequality fails for such process and the measure on the underlying state space does not satisfy volume doubling property.
May 1, 2013
Prof. Peng Wang :
4 p.m. in SEO 636
Abstract
In longitudinal studies, mixed-effects models are important for addressing
subject-specific effects. However, most existing approaches
assume a normal distribution for the random effects, and this could affect the bias and efficiency of the fixed-effects estimator.
Even in cases where the estimation of the fixed effects is robust with a misspecified distribution of the random effects, the
estimation of the random effects could be invalid. We propose a new approach to estimate
fixed and random effects using conditional quadratic inference
functions. The new approach does not require the specification of
likelihood functions or a normality assumption for random effects. It can
also accommodate serial correlation between observations within the same cluster,
in addition to mixed-effects
modeling. Other advantages include not
requiring the estimation of the unknown variance components associated
with the random effects, or the nuisance parameters associated with the working
correlations. We establish asymptotic results for the fixed-effect parameter estimators which do not rely on the consistency of the random-effect estimators.
Some applications of the proposed approach will also be presented.
Aug. 28, 2013
Statistics group :
4 p.m. in SEO 636
Abstract
Welcome new students and discuss and expose students to various statistical associations. This gathering is partially sponsored by the American Statistical Association.
Sept. 4, 2013
Dr. Lawrence Lin :
4 p.m. in SEO 636
Abstract
This will be a general overview presentation with practical examples and without much statistical formulas. We will introduce the concepts of un-scaled and scaled agreement statistics based on the basic case between two raters with paired samples for continuous, binary, and ordinal data. We will then progress into more complex cases when we have multiple raters and each rater has multiple readings per sample. Here, we can assess intra-rater and inter-rater agreement, compare inter-rater deviation to intra-rater deviation, and compare precision of a rater against another. We will explore the meaning of the two-stage criteria presented in the FDA guidance UCM070244: Statistical Approaches to Establishing Bioequivalence. The content is largely based on the materials presented in the newly published book by Springer, entitled "Statistical Tools for Assessing Agreement".
Sept. 11, 2013
Dr. Caiyan Li :
4 p.m. in SEO 636
Abstract
Graphs and networks are common ways of depicting biological information.
In biology, many different biological processes are represented by graphs,
such as regulatory networks, metabolic pathways and protein protein interaction networks.
This kind of a priori use of graphs is a useful supplement to the standard numerical
data such as microarray gene expression data. In this presentation, we consider the
problem of regression analysis and variable selection when the covariates are linked
on a graph. We study a graph constrained regularization procedure and its theoretical
properties for regression analysis to take into account the neighborhood information of
the variables measured on a graph. This procedure involves a smoothness penalty on the
coefficients that is defined as a quadratic form of the Laplacian matrix
associated with the graph. We establish estimation and model selection
consistency results and provide estimation bounds for both fixed and
diverging numbers of parameters in regression models. We also developed a
second method using Markov Random Field to incorporate the graph information
into analysis of high dimensional data. Finally, we demonstrate by simulations
and a real dataset that the proposed procedure can lead to better variable selection
and prediction than existing methods that ignore the graph information associated
with the covariates.
Sept. 18, 2013
Dr. Stefanie Biedermann :
4 p.m. in SEO 636
Abstract
Finding optimal designs for nonlinear models is challenging in general. Although some recent results allow us to focus on a simple subclass of designs for most problems, deriving a specific optimal design still mainly depends on numerical approaches. There is need for a general and efficient algorithm which is more broadly applicable than the current state of the art methods. We present a new algorithm which can be used to find optimal designs with respect to a broad class of optimality criteria, when the model parameters or functions thereof are of interest, and for both locally optimal and multi-stage design strategies. We prove convergence to the optimal design, and show in various examples that the new algorithm outperforms the current state of the art algorithms.
Sept. 25, 2013
Prof. Jian (Frank) Zou :
4 p.m. in SEO 636
Abstract
The complexity of spatio-temporal data in epidemiology and surveillance presents challenges such as low signal-to-noise ratio and generating high false positive rate for researchers and public health agencies. Central to the problem in the context of disease outbreaks is a decision structure that requires trading off false positives for delayed detections. We describe a novel Bayesian hierarchical model capturing the spatio-temporal dynamics in public health surveillance data sets. We further quantify the performance of the method to detect outbreaks by incorporating different criteria, including false alarm rate, timeliness and cost functions. Our data set is derived from emergency department (ED) visits for Influenza-like illness and respiratory illness in the Indiana Public Health Emergency Surveillance System (PHESS). The methodology incorporates Gaussian Markov random field (GMRF) and spatio-temporal conditional autoregressive (CAR) modeling. Features of this model include timely detection of outbreaks, robust inference to model misspecification, reasonable prediction performance, as well as attractive analytical and visualization tool to assist public health authorities in risk assessment. Our numerical results show that the model captures salient spatio-temporal dynamics that are present in public health surveillance data sets, and that it appears to detect both "annual" and "atypical" outbreaks in a timely, accurate manner. We present maps that help make model output accessible and comprehensible to public health authorities. We use an illustrative family of decision rules to show how output from the model can be used to inform false positive--delayed detection tradeoffs.
Oct. 2, 2013
Jan Hannig :
4 p.m. in SEO 636
Abstract
R. A. Fisher's fiducial inference has been the subject of many discussions
and controversies ever since he introduced the idea during the 1930's. The idea experienced a bumpy ride,
to say the least, during its early years and one can safely say that it eventually fell into disfavor among
mainstream statisticians. However, it appears to have made a resurgence recently under various names and modifications.
For example under the new name generalized inference fiducial inference has proved to be a
useful tool for deriving statistical procedures for problems where
frequentist methods with good properties were previously unavailable.
Therefore we believe that the fiducial argument of R.A. Fisher deserves a
fresh look from a new angle.
In this talk we first generalize Fisher's fiducial argument and obtain a
fiducial recipe applicable in virtually any situation. We demonstrate this
fiducial recipe on examples of varying complexity. We also investigate, by
simulation and by theoretical considerations, some properties of the
statistical procedures derived by the fiducial recipe showing they often
posses good repeated sampling, frequentist properties.
Finally, we show how a generalized fiducial inference paradigm can be used to combine
confidence distributions.
Portions of this talk are based on a joined work with Hari Iyer, Thomas
C.M. Lee, Randy Lai and Min-ge Xie.
Oct. 9, 2013
Ryan Martin :
4 p.m. in SEO 636
Abstract
Estimating a sparse high-dimensional normal mean vector is an
important classical problem. In this talk, I will introduce a new
empirical Bayes model based on a unique data-dependent prior. I will
show that, under some conditions, our empirical Bayes posterior
distribution concentrates on balls, centered at the true mean vector,
with squared radius proportional to the frequentist minimax rate for the
given sparsity class. This result provides some new insight concerning
the fully Bayes approach to this same problem. Asymptotic minimaxity of
the corresponding empirical Bayes posterior mean is shown, and a simple
Gibbs sampling algorithm for computation will be discussed. Finally,
two simulation studies will be presented, demonstrating the strong
finite-sample performance of the proposed estimator against a variety of
popular alternatives. (This is joint work with Stephen Walker at the
University of Texas at Austin.)
Oct. 16, 2013
Prof. Ching-Shui Cheng :
4 p.m. in SEO 636
Abstract
Design key is a useful method of constructing factorial designs with what J. A. Nelder called simple block structures. One familiar method of constructing such designs is to use independent treatment factorial effects to divide the treatment combinations into blocks, rows, and columns, etc. Certain conditions need to be verified to ensure that the desired block structure is achieved. The method of design key, proposed by H. D. Patterson, on the other hand, guarantees the desired block structure by choosing some appropriate contrasts of the experimental units to be aliases of the main-effect contrasts of treatment factors. We provide some useful templates for implementing design key construction. Some constraints imposed by the block structure are built into the template. This eliminates the need to check some conditions for design eligibility.
Oct. 23, 2013
Prof. Yaozhong Hu :
4 p.m. in SEO 636
Abstract
The classical central limit theorem is one of the most important
theorem in probability theory. The theorem states that if $X_1$, $\cdots$,
$X_n$ are independent identically distributed random variables and if $F_n$
is the difference between the sample mean and the mean of the random
variables properly normalized, then $F_n$ converges to a normal
distribution in
distribution. Recent results extend this results to other random variables
for example given by Wiener chaos (multiple It\^o-Wiener integrals).
In this talk, we shall obtain some conditions on $F_n$ such that the
distributions of the random variables $F_n$ have densities $f_n(x)$ with
respect to Lebesgue measure and $f_n(x)$ converges to the normal density
$\phi(x)=\frac{1}{\sqrt{2\pi}}e^{-|x|^2/2}$.
The tool that we use is the Malliavin calculus and a brief introduction
will also be given. This is an ongoing joint work with Fei Lu and David
Nualart
Oct. 30, 2013
Ms. Ping Jiang :
4 p.m. in SEO 636
Abstract
Working as a statistician in a pharmaceutical company is exciting and rewarding. This presentation will show the vital roles played by statisticians in various stages of drug research and development. The information is intended for the students who are interested in pursuing a career in pharmaceutical industry.
Nov. 6, 2013
Prof. Weixing Song :
4 p.m. in SEO 636
Abstract
An estimation procedure is proposed for a nonparametric regression in which some covariates are measured with errors and some are not. The procedure combines the ordinary and deconvolution kernel estimation techniques. It is shown that the optimal local and global convergence rates, as well as the uniform convergence rate over a class of joint distributions of the response and the covariates, depend on the tail behavior of the characteristic functions of the measurement error distributions. Examples are given to show the general applicability of the proposed methodology, and finite sample performance is evaluated by some numerical simulation studies.
Nov. 13, 2013
Prof. Qiongxia Song :
4 p.m. in SEO 636
Abstract
For time series nonparametric regression models with discontinuities, we propose to use polynomial splines to estimate locations and sizes of jumps in the mean function. Under reasonable conditions, test statistics for the existence of jumps are given and their limiting distributions are derived under the null hypothesis that the mean function is smooth. Simulations are provided to check the powers of the tests. A climate data application and an application to the U.S. unemployment rates of men and women are used to illustrate the performance of the proposed method in practice.
Nov. 20, 2013
Prof. Harry Crane :
4 p.m. in SEO 636
Abstract
In fields as diverse as physics, biology, sociology and national security,
complex networks are used to model structural relationships among individuals
and variables. In many applications, the networks vary over time and so are
appropriately modeled by a stochastic process on the space of graphs.
Motivated by these applications, we consider Markov processes that evolve
on the space of infinite graphs. Natural statistical models for such processes
are both exchangeable with respect to relabeling vertices and have the property
that all restrictions to finite induced subgraphs are finite state space Markov
chains. Our main theorem provides a Levy-Ito-type characterization for all
processes in this class. Our approach also gives a straightforward recipe for
simulating general processes of this type, which may be useful in a range of
applications.
Jan. 22, 2014
Shuwen Lou :
4 p.m. in SEO 636
Abstract
Think of installing an infinite pole on top of the ground. We want to model the random movement of an ant on this space. However, as we know, the standard 2-dimensional Brownian motion does not hit a single, which means that once the ant is on the ground, it will never have the chance to climb up the pole. We fix this problem by defining Brownian motion with varying dimension on this state space as a darning process whose rigorous definition will be introduced in the talk. The main results are about global two-sided heat kernel estimates for such processes. We will see from the heat kernel estimates that these processes embody both 1-dimensional property and 2-dimensional property, which depends not only on the regions of the points but also on time.
Jan. 29, 2014
Prof. Kunpeng Zhang :
4 p.m. in SEO 636
Abstract
Social media has become a popular platform that connects people who share information, in particular personal opinions. Through such a fast information exchange mechanism, reputation of individuals, consumer products, or business companies can be quickly built up within a social network. Recently, applications mining social network data start emerging to find the communities sharing the same interests for marketing purposes. Knowing the reputation of social network entities, such as celebrities or business companies, can help develop better strategies for election campaigns or new product advertisements. In this work, we propose a probabilistic graphical model to collectively measure reputations of entities in social networks. By collecting and analyzing large amount of user activities on Facebook, our model can effectively and efficiently rank entities, such as presidential candidates, professional sport teams, musician bands, and companies, based on their social reputation. The proposed model produces results largely consistent with the two publicly available systems - movie ranking in Internet Movie Database and business school ranking by the US news & World Report - with the correlation coefficients of 0.75 and −0.71, respectively. In addition, I will briefly talk about other projects I am working on: (1) sentiment identification of social media data, and (2) finding target users for online brand advertising based on a large amount of user historical activities.
Feb. 12, 2014
Prof. Yazhen Wang :
4 p.m. in SEO 636
Abstract
Quantum computation and quantum information are of great current interest in computer science, mathematics, physical sciences and engineering. They will likely lead to a new wave of technological innovations in communication, computation and cryptography. As the theory of quantum physics is fundamentally stochastic, randomness and uncertainty are deeply rooted in quantum computation, quantum simulation and quantum information. Consequently quantum algorithms are random in nature, and quantum simulation utilizes Monte Carlo techniques extensively. Thus statistics can play an important role in quantum computation and quantum simulation, which in turn offer great potential to revolutionize computational statistics. This talk will give a brief review on quantum computation, quantum simulation and quantum information. I will first introduce the basic concepts of quantum computation and quantum simulation and then present my recent work on statistical analysis of quantum systems with applications to quantum computation and quantum simulation.
March 12, 2014
Prof. Annie Qu :
4 p.m. in SEO 636
Abstract
We develop new modeling and estimation for personalized treatment for individuals with high heterogeneity. Incorporating subject-specific information into treatment subgroup is critical since individuals could react to the same treatment quite differently. We propose to identify subgroups with longitudinal observations through random-effects estimation where the random effects are not necessarily normal distributed. The advantage of this approach is that we can quantify intrinsic associations between unobserved subject-specific effects and observed treatment outcomes, and therefore provide optimal treatment assignments for different individuals. In contrast, traditional mixed-effects models assuming normal distribution cannot effectively distinguish different patterns of treatment effects. We develop asymptotic consistency theory for individual treatment effect estimation, and show that the new estimator is more efficient than the random effect estimator which ignores correlation information from longitudinal data. Simulation studies and a data example from an AIDS clinical trial group confirm that the proposed method is quite efficient in identifying an effective treatment strategy for subgroups in finite samples. This is joint work with Hyunkeun Cho and Peng Wang.
March 19, 2014
Dr. Hongwen Guo :
4 p.m. in SEO 636
Abstract
Under formula score instruction (FSI), test takers may omit items instead of guessing. If the students are required to answer every items (under the rights only scoring instruction, ROI), the score distribution will be different. In this study, we tried to model the data using a simple statistical model and provide a formula to predict the score distribution under ROI based on the score distribution and the omit rate observed under FSI. A preliminary investigation of the guessing parameter is presented in the paper. Based on the data used in the study, the guessing parameter may be close or slightly lower than the chance score.
April 2, 2014
Prof. Bikas Sinha :
4 p.m. in SEO 636
Abstract
Noted environmental statistician Professor Anil Gore, University
of Pune, dealt with the problem of estimation of non-identifiable bird
species in BNHS Survey. This study was conducted around mid 1970's. In
recent time, Dr. Tommy Wright of US Census Bureau posed a similar problem.
The speaker had the opportunity to be involved in both the studies in a
limited way. He will discuss salient features of the problem and the
proposed solutions.
April 9, 2014
Paul Livermore Auer :
4 p.m. in SEO 636
Abstract
To date, much of the work in understanding the genetic basis of common diseases has focused on genetic variants that exist with high frequency in human populations. Recently, the focus has shifted to investigating the role that rare genetic variants may play. Genetic studies of rare variants bring with them a host of statistical issues including decreased statistical power and missing data. We examine these issues and demonstrate the performance of solutions with both real and simulated data.
April 16, 2014
Prof. Ming-Hung (Jason) Kao :
4 p.m. in SEO 636
Abstract
Functional magnetic resonance imaging (fMRI) is one of the leading brain mapping technologies for studying brain activity in response to mental stimuli. For neuroimaging studies utilizing this pioneering technology, there is a great demand of high-quality experimental designs that help to collect informative data to make precise and valid inference about brain functions. In this talk, I briefly introduce some recently developed analytical and computational results on fMRI experimental designs. The performance of some commonly considered designs such as m-sequences is discussed. In addition, a new type of fMRI designs that are constructed using a certain type of Hadamard matrices is also discussed. Under certain assumptions, these designs can be shown to be optimal in some statistically meaningful sense. Some possible future research directions are also presented.
April 30, 2014
Prof. Dacheng Xiu :
4 p.m. in SEO 636
Abstract
Estimating the covariance between assets using high frequency data is challenging due to market
microstructure effects and asynchronous trading. In this paper we develop a multivariate realised
quasi-likelihood (QML) approach, carrying out inference as if the observations arise from an asynchronously
observed vector scaled Brownian model observed with error. Under stochastic volatility
the resulting realised QML estimator is positive semi-definite, uses all available data, is consistent and
asymptotically mixed normal. The quasi-likelihood is computed using a Kalman filter and optimised
using a relatively simple EM algorithm which scales well with the number of assets. We derive the
theoretical properties of the estimator and prove that it achieves the efficient rate of convergence. The
estimator is also analysed using Monte Carlo methods and applied to equity data with varying levels
of liquidity.
Aug. 13, 2014
Abhyuday Mandal :
11 a.m. in SEO 612
Abstract
Random effects models play an important role in model-based
small area
estimation. Random effects account for any lack of fit of a regression
model for the population means of small areas on a set of explanatory
variables. In a recent paper, Datta, Hall and Mandal (2011, J. Amer.
Statist. Assoc.) showed that if the random effects to account for a lack
of fit of a regression model can be dispensed with through a statistical
test, then the model parameters and the small area means can be estimated
with substantially higher accuracy. The work of Datta et al. (2011) is
most useful when the number of small areas, m, is moderately large. For
large m, the null hypothesis of no random effects will likely be rejected.
Rejection of null hypothesis is usually caused by a few large residuals
signifying a departure of the direct estimator (Yi) from the synthetic
regression estimator. As a flexible alternative to the Fay-Herriot random
effects model and the approach in Datta et al. (2011), in this paper we
consider a mixture model for random effects. It is reasonably expected
that small areas with population means explained adequately by covariates
have little model error, and the other areas with means not adequately
explained by covariates will require a random component added to the
regression model. This model is a flexible alternative to the usual random
effects model and the data determine the extent of lack of fit of the
regression model for a particular small area, and include a random effect
if needed. Unlike the Datta et al. (2011) approach which recommends
excluding random effects from all small areas if a test of null hypothesis
of no random effects is not rejected, the present model is less
restrictive. We used this mixture model to estimate poverty ratios for 5-
to 17-year old related children for the 50 U.S. states and Washington, DC.
This application is motivated by the SAIPE project of the US Census
Bureau. We empirically evaluated the accuracy of the direct estimates and
the estimates obtained from our mixture model and the Fay-Herriot random
effects model. These empirical evaluations and a simulation study, in
conjunction with a measure of uncertainty of the new estimates show that
they are more accurate than the frequentist and the Bayes estimates
resulting from the standard Fay-Herriot model.
Aug. 27, 2014
Statistics Faculty :
4 p.m. in SEO 636
Sept. 3, 2014
Keli Liu :
4 p.m. in SEO 636
Abstract
Priors are the path to the dark side. Fisher developed the Fiducial argument to obtain prior free "posterior" inferences but at the
seeming cost of violating basic probability laws. Was Fisher crazy or did madness mask innovation? Fiducial calculations can be
easily understood through the missing-data perspective which illuminates for us that the Fiducial "posterior" is in fact a prior updated
not with the full data likelihood, but a <i>partial</i> likelihood in the spirit of Cox regression. Just as Cox regression arose from a
need to render inferences robust to an unknown hazard function, so Fiducial inferences are insensitive to the prior. While Statistics has
fixated two extremes---fully conditional (but fragile) Bayesian inferences or unconditional (but robust) Frequentist inferences---a
compromise via partial conditioning has gone ignored. Surely, the middle ground is more fiducial than either extreme.
Sept. 17, 2014
Bikas Sinha :
4 p.m. in SEO 636
Abstract
Considered is the problem of estimation of the size of a finite labeled population. However, the units of the population are not directly accessible. This is a well defined finite population of Reference Units (RUs) and these RUs may be accessed directly. As against this, the former population is referred to as that of Ultimate Units (UUs). Further to this, there is a well defined network connecting the population of RUs to that of the UUs. A sample of RUs will naturally create a partial view of the population network. Based on this sample network, it is required to unbiasedly estimate the size of the population of UUs. We review the literature (which is scanty anyway) and provide some suggestions towards satisfactory acceptable solutions to this fascinating problem.
Sept. 24, 2014
Jennifer Pajda-Delao :
4 p.m. in SEO 636
Abstract
I will give some interesting estimates of Brownian motion on
manifolds, including exit time estimates from a geodesic ball.
Oct. 1, 2014
Ruoting Gong :
4 p.m. in SEO 636
Abstract
The small-time asymptotic behavior of option prices and implied volatilities for jump-diffusion models has received much attention in recent years. In this presentation, we study the time-to-maturity asymptotics of call option prices under a variety of models with Lévy jumps. In the out-of-the-money (OTM) and in-the-money (ITM) case, we consider a general stochastic volatility model with independent Lévy jumps for the log-return process of the underlying stock price. In this setting, small-time expansions, of arbitrary polynomial order, in time-t, are obtained for both OTM and ITM call option prices. In the at-the-money (ATM) case, a novel second-order approximation of the call option price is obtained for a large class of exponential "tempered-stable-like" Lévy models with or without Brownian component. As a consequence, small-time expansions of the corresponding Black-Scholes implied volatilities are also addressed in both cases. This is the joint work with J. E. Figueroa-López and C. Houdré.
Oct. 8, 2014
Stat Grad Students :
4 p.m. in SEO 636
Oct. 15, 2014
Mengyu Xu :
4 p.m. in SEO 636
Abstract
We develop an asymptotic theory for $L^2$ norms of sample mean vectors of high-dimensional data. An invariance principle for the $L^2$ norms is derived under conditions that involve a delicate interplay between the dimension $p$, the sample size $n$, and the moment condition. Under proper normalization, central and non-central limit theorems are obtained. To facilitate the related statistical inference, we propose a resampling calibration method to approximate the distributions of the $L^2$ norms. Our results are applied to multiple tests and inference of covariance matrix structures.
Oct. 22, 2014
Wei Zheng :
4 p.m. in SEO 636
Abstract
A systematic study is carried out regarding universally optimal designs under the interference model, previously investigated by Kunert and Martin (2000) and Kunert and Mersmann (2011). Parallel results are also provided for the undirectional interference model, where the left and right neighbor effects are equal. It is further shown that the efficiency of any design under the latter model is at least its efficiency under the former model. Designs universally optimal for both models are also identified. Most importantly, this paper provides Kushner's type linear equations system as a necessary and sufficient condition for a design to be universally optimal. This result is novel for models with at least two sets of treatment-related nuisance parameters, which are left and right neighbor effects here. It sheds light on other models in deriving asymmetric optimal or efficient designs.
Oct. 29, 2014
V. Devanarayan :
4 p.m. in SEO 636
Abstract
Biomarkers such as those based on genomic, proteomic and imaging
modalities play a vital role in biopharmaceutical R&D. Examples include
the discovery of novel genes/targets related to various diseases based on
which a suitable therapeutic can be developed, diagnostics for different
disease subtypes, identification of patients that are more likely to
progress in disease or benefit from a particular therapeutic, etc. The
discovery of such biomarkers are typically based on the evaluation of
high-dimensional datasets that require a strong combination of
bioinformatic and statistical considerations. This seminar will provide
a practical overview and intuitive explanation of some important concepts
and considerations around the analyses of such high-dimensional data.
Nov. 5, 2014
Sonja Petrovic :
4 p.m. in SEO 636
Abstract
The ubiquity of network data in the world around us does not imply that the
statistical modeling and fitting techniques have been able to catch up with
the demand. This talk will discuss some of the basic modeling questions
that every statistician knows are fundamental, some of the recent advances
toward answering them, and the challenges that remain. The specific focus of the talk will be on goodness of fit testing for
random graph models. Recent joint work with Despina Stasi and Elizabeth
Gross developed a new testing framework for graphs that is based on
combinatorics of hypergraphs and model geometry. I will summarize our work
by showing simulation results for the popular $p_1$ model for directed
random graphs.
Nov. 12, 2014
Xianggui Qu :
4 p.m. in SEO 636
Abstract
High-throughput screening (HTS) is a large-scale process that screens hundreds of thousands to millions of compounds in order to identify potentially leading candidates rapidly and accurately. There are many statistically challenging issues in HTS. In this talk, I will focus the spatial effect in primary HTS. I will discuss the consequences of spatial effects in selecting leading compounds and why the current experimental design fails to eliminate these spatial effects. A new class of designs will be proposed for elimination of spatial effects. The new designs have the advantages such as all compounds are comparable within each microplate in spite of the existence of spatial effects; the maximum number of compounds in each microplate is attained, etc. Optimal designs are recommended for HTS experiments with multiple controls.
Nov. 19, 2014
Raymond Mess / Nick Syring :
4 p.m. in SEO 636
Abstract
This is a special graduate student-organized seminar in which two PhD students (Raymond Mess and Nick Syring) will give 20+
minute talks about their ongoing research. The respective abstracts are below.
(Mess) In this talk, I will introduce the new double empirical Bayes framework, which is based on the use of data to both center and
regularize the prior. An application of this framework to the problem of inference in the sparse (p >> n) linear model will also
be presented.
(Syring) I will introduce a method to obtain Bayesian-like posterior inference
for an unknown parameter without the need for a likelihood. Such a
method makes producing interval estimates straightforward while avoiding
problems that may arise from model misspecification. Finally, I will
discuss an application of this approach to an important problem in medical statistics.
Dec. 3, 2014
Dr. Devon Lin :
4 p.m. in SEO 636
Abstract
Computer experiments with qualitative and quantitative factors occur frequently in various applications in science and engineering. Design and analysis of such experiments is not yet completely resolved. To address this issue, we propose a new class of designs and a flexible modeling approach. The proposed designs allow us to accommodate a large number of qualitative factors with economic run sizes. Properties of such designs will be discussed. Several construction methods will be given. The new modeling approach employs a flexible function to capture the correlation among qualitative and quantitative factors. Several examples are provided to demonstrate significant improvement in prediction.
Jan. 14, 2015
Zhenan Wang :
4 p.m. in SEO 636
Abstract
We will start on the classical De Giorgi iteration for parabolic PDEs. We will explain how a stochastic version of De Giorgi iteration can be developed and applied to prove H\"older continuity for solution of
stochastic partial differential equations with measurable coefficient. We will also introduce fine properties for the solutions obtained by applying the stochastic De Giorgi iterations.
Jan. 21, 2015
Ryan Martin :
4 p.m. in SEO 636
Abstract
The Bayesian framework provides a nice recipe for constructing the predictive distribution of a future observation given the available data.
Except for simple problems, computation of the Bayes predictive requires Monte Carlo which cannot be done recursively. However, when data is received sequentially, e.g., in finance applications, a recursive update to the predictive distribution is desired. In this talk, I will explain how
the Bayes predictive step can be rewritten using a copula, which makes recursive updates of the predictive distribution possible. This new
representation motivates a version of Newton's predictive recursion algorithm for the predictive density, which can be used for fast and
universal recursive predictive density estimation. Illustrations and convergence theory for the new algorithm is provided.
Jan. 28, 2015
Janna Lierl :
4 p.m. in SEO 636
Abstract
I will present sharp two-sided bounds for the Dirichlet heat kernel on bounded domains. The domain is assumed to satisfy an inner uniformity condition. For example, the interior of the Koch snowflake is an inner uniform domain,
as is any convex domain, or the complement of any convex domain in Euclidean space.
More generally, we have considered the Dirichlet heat kernel on domains in a metric measure Dirichlet space, assuming the space satisfies a Poincar\'e inquality and has the volume doubling property. We have also considered non-symmetric Dirichlet spaces. In particular, we can estimate the Dirichlet heat kernel if the kernel is associated with a differential operator in divergence form with bounded measurable coefficients and symmetric uniformly elliptic second order part. This talk is based on a joint paper with Laurent Saloff-Coste.
Feb. 11, 2015
Qianshun Cheng / Tian Tian :
4 p.m. in SEO 636
Abstract
(Cheng) An experiment often has several competing objectives cannot be characterized by only one of the standard optimality criteria.
Multiple objective optimal design aims to optimize the target objective while guarantee that efficiency of the other objectives interested
are above acceptable levels. Such optimality problem is in general challenging and typically be solved through algorithm approach.
The existing approaches either have high computation cost or have low accuracy. In this talk, I will present a new algorithm which can
be used for general multiple objective optimal design problems regardless of model settings. Compared with the existing approach, the
new algorithm enjoys low computation cost and high accuracy.
(Tian) A widely used approach of designing the phase I clinical trial is continual reassessment method (CRM), which has been shown
through many simulations to be more effective than other traditional approaches. In this talk, I will show that the CRM algorithm is
indeed efficient from the perspective of optimal design theory. Specifically, simple power model and logistic model -- two popular
models, are considered. For simple power model, I'll show the efficiency of CRM depends on the target toxicity rate and CRM is highly
efficient in practice. A remarkable fact is that the optimal design selects the dose level such that the corresponding toxicity rate is
around 0.2, which is exactly the commonly used target toxicity rate in clinical trials. Moreover, by incorporating the idea of optimal
design into the study, the percentage of toxicity occurrence in the trial will drop by a great amount. As for logistic model, I'll show that
the CRM approach is indeed optimal, which will justify the efficiency of the algorithm in theory.
Feb. 18, 2015
Haiying Wang :
4 p.m. in SEO 636
Abstract
For massive data with super-large sample size n, it is computationally infeasible
to obtain maximum likelihood estimates for unknown parameters, especially when the
estimator does not have a close-form solution. This paper proposes fast leveraging
algorithms to efficiently approximate the maximum likelihood estimates of unknown
parameters in logistic regression models with binary responses, one of the most commonly
used models in practice for classification. We theoretically prove the consistency
of the leveraging algorithms, develop nearly optimal two-step leveraging strategies, and
evaluate the performance of the proposed methods using synthetic and real data sets.
Feb. 25, 2015
Wei-kuo Chen :
4 p.m. in SEO 636
Abstract
Spin glasses are disordered spin systems originated from the desire of understanding the strange magnetic behaviors of certain alloys in physics. As mathematical objects, they are often cited as examples of complex systems and have provided several fascinating structures and conjectures. This talk will be focused on one of the most famous mean-field spin glasses, the Sherrington-Kirkpatrick model. We will present results on the conjectured properties of the Parisi measure including its uniqueness and quantitative behaviors. This is based on joint works with A. Auffinger.
March 4, 2015
Lei Yang :
4 p.m. in SEO 636
Abstract
Variable selection is popular in high-dimensional data analysis to identify the truly informative variables. Many variable selection methods have been developed under various model assumptions, such as linear model and additive model. However, their success largely rely on validity of the assumed models. In this talk, I will introduce a model-free variable selection method based on gradient learning. The key idea is that if a variable is informative is equivalent to if its corresponding gradient function is substantially non-zero. The proposed method is formulated in a framework of learning gradients equipped with a flexible reproducing kernel Hilbert space. Computationally, a blockwise majorization decent (BMD) algorithm is introduced for efficient computation. Theoretically, without assuming explicit models, the estimation and variable selection consistencies are established. A variety of simulated examples and real-life examples are provided to evaluate the performance.
March 18, 2015
Professoe Suojin Wang :
4 p.m. in SEO 636
Abstract
In this talk, we present a new double order selection test for checking second-order stationarity of a time series. To develop the test, a sequence of systematic samples are defined via the Walsh functions. Then the deviations of the autocovariances based on these systematic samples from the corresponding autocovariances of the whole time series are calculated and the uniform asymptotic joint normality of these deviations over different systematic samples is obtained. With a double order selection scheme, our test statistic is constructed by combining the deviations at different lags in the systematic samples. The null asymptotic distribution of the proposed statistic is derived and the consistency of the test is shown under fixed and local alternatives. Simulation studies demonstrate well-behaved finite sample properties of the proposed method. Comparisons with some existing tests in terms of power are given both analytically and empirically. In addition, the proposed method is applied to check the stationarity assumption of a chemical process viscosity readings data.
April 8, 2015
Tiefeng Jiang :
4 p.m. in SEO 636
Abstract
We study the eigenvalues of a Laplace-Beltrami operator defined on the set of the symmetric polynomials, where the eigenvalues are expressed in terms of partitions of integers. By assigning partitions with the uniform measure, the restricted uniform measure, the Plancherel measure, or the restricted Jack measure, we prove that the global distributions of the eigenvalues are asymptotically the Gumbel distribution, a new distribution F, the Tracy-Widom distribution and the Gamma distribution, respectively. An explicit representation of F is obtained by a function of independent random variables. We also derive an independent result on random partitions itself: a law of large numbers for the restricted uniform measure. This is a joint work with Ke Wang
April 15, 2015
Xin Huang :
4 p.m. in SEO 636
Abstract
Mechanistic relationships between the clinical outcome (efficacy or safety
endpoints) versus putative biomarkers, clinical baseline and related
predictors are usually unknown, and must be deduced empirically from
experimental data. Such relationships enable the implementation of a
personalized medicine strategy in clinical trials to help stratify patients
in terms of disease progression, clinical response, treatment differentiation,
etc. These relationships are often requires complex modelling to develop the
prognostic and predictive signatures. For the purpose of easier interpretation
and implementation in the clinical practice, defining a multivariate biomarker
signature in terms of thresholds on the biomarker combinations is preferable.
In this presentation, we propose some methods for developing such signatures
in the context of continuous, binary and time-to-event endpoints. Results from
simulations and case-study illustration will also be provided.
April 22, 2015
John Hardwick / Brian Powers :
4 p.m. in SEO 636
Abstract
(Hardwick) An assignment game is a cooperative game with its player set divided into
two groups, where a coalition receives a positive payoff if and only if it
contains at least one player from both groups. In other words, it concerns
a bipartite matching, with a payoff defined for each pair (payoffs for
larger coalitions are determined additively). The 1994 paper by Solymosi
and Raghavan gives an algorithm for finding the nucleolus (an optimal
solution) of such a game, requiring O(n4) operations (assuming that n is
the size of both groups). In this paper, we examine the special case in
which the payoffs for each pair take only binary values, which we can
think of as an indicator of whether or not the pair is compatible. For
this case, we present a new algorithm which capitalizes on the graphical
aspect of the previous one. This algorithm is shown to be an improvement,
requiring only O(n3) operations. We also discuss sociological implications
of this solution, considering each pair’s split of the payoff to be their
relationship’s balance of power.
(Powers) Final-Offer Arbitration is used in major league baseball to determine
players' salaries, and has been used in various other wage disputes. In
Final-Offer Arbitration, two contesting parties each provide a mediator
with a final offer. The mediator must choose one of the two offers with no
option for compromise. Under certain conditions, Nash equilibria are known
for the single-variable case. Here we will explore solution points when
players have multiple components to their offer, effectively bringing the
dimension of the game to 2 or higher.
April 29, 2015
Stat faculty and students :
4 p.m. in SEO 636
Abstract
This panel discussion will give students an opportunity to have questions about various things (e.g., advantages and disadvantages of
an academic career, general strategies for successful research, career opportunities outside of academia and how to prepare for them, etc)
answered by faculty and senior graduate students.
Sept. 2, 2015
Statistics Faculty & Students :
4 p.m. in SEO 636
Sept. 9, 2015
Nick Syring :
4 p.m. in SEO 636
Abstract
In some applications, the relationship between the observable data and unknown parameters is described via a loss function rather than
likelihood. In such cases, the standard Bayesian methodology cannot be used, but a Gibbs posterior distribution can be constructed by
appropriately using the loss in place of a likelihood. Inference based on the Gibbs posterior is not straightforward, however,
because the finite-sample performance is highly sensitive to the scale of the loss function. In this talk, I will propose a Gibbs
Posterior Scaling (GPS) algorithm that adaptively selects the scaling in order to calibrate the corresponding Gibbs posterior credible
regions. Two examples, namely, classification and quantile regression, are used to demonstrate that the Gibbs posterior with scale chosen
by GPS produces valid interval estimates which are at least as efficient those obtained from other methods.
Sept. 16, 2015
Stat Faculty and Students :
4 p.m. in SEO 636
Sept. 23, 2015
Daniel Conus :
4 p.m. in SEO 636
Abstract
The talk will focus on results related to the notion of intermittency: i.e. the property that a random field develops large values ("high peaks") when time gets large. We will first present it with known examples and, then, describe how this phenomenon appears in the context of SPDEs. In particular, we will illustrate how the intermittent behavior of the solution to an SPDE depends on the type of driving noise in the case of stochastic heat and wave equations driven by fractional noise. In the latter case, the results are obtained via a Feynman-Kac representation of the moments similar to the one introduced in Dalang-Mueller-Tribe (2008). This is based on a joint work with Raluca Balan (Univ. of Ottawa).
Sept. 30, 2015
Hsin-Hsiung Huang :
4 p.m. in SEO 636
Abstract
The Natural Vector combined with Hausdorff distance has been successfully applied for classifying and clustering multiple-segmented viruses. Additionally, k-mer
methods also yield promising results for global genome comparison. It is not known whether combining these two approaches can lead to more accurate results. The
author proposes a method of combining the Hausdorff distances of the 5-mer counting vectors and natural vectors which achieves the best classification without
cutting off any sample. Using the proposed method to predict the taxonomic labels for the 2,363 NCBI reference viral genomes dataset, the accuracy rates are
96.95%, 94.37%, 99.41% and 93.82% for the Baltimore, family, subfamily, and genus labels, respectively. We further applied the proposed method to 48 isolates of
the influenza A H7N9 viruses which have eight complete segments of nucleotide sequences. The single-linkage clustering trees and the statistical hypothesis testing
results all indicate that the proposed ensemble distance measure can cluster viruses well using all of their segments of genome sequences.
Oct. 7, 2015
Naitee Ting :
4 p.m. in SEO 636
Abstract
In the process of drug discovery and drug development, understanding the dose-response relationship is one of the most challenging tasks. It is also critical to identify the right range of doses in early stages of clinical development so that Phase III trials can be designed to confirm these doses. Usually at the beginning of Phase II, there is not a lot of available information to help guiding the study design. At this stage, Phase II clinical studies are needed to establish proof of concept (PoC), to identify a set of potentially effective and safe doses, and to estimate dose-response relationships.
Challenges in designing these studies include: selection of the dose frequency and the dose range, choice of clinical endpoints or biomarkers, and use of control(s), among others. Consequences of bad Phase II study designs may lead to the delay of the entire clinical development program or the waste of R&D investment. Misleading results obtained from poor designs could cause a Phase III program to confirm a wrong set of doses, or to stop developing a potentially useful drug. Therefore, it is critical to consider an entire drug development plan, to make best use of all the available information, and to include all relevant experts in designing Phase II dose response clinical trials. This presentation discusses some of these considerations.
Oct. 14, 2015
Jonathan Stallings :
4 p.m. in SEO 636
Abstract
Standard design criteria like the A-, E-, and D-criterion implicitly assume the experimenter is equally interested in all estimable functions. Because of this, efficient designs under these criteria spread information evenly across the estimation space. In some cases, optimal designs can be analytically derived under these criteria but researchers are beginning to rely on design search algorithms to find these designs. These computer-generated designs are often found under the D-criterion because of its fast computations with point- and coordinate-exchange algorithms. However, the D-criterion is a poor assessment of a design when the goals of an experiment imply differential interest among the estimable functions. To reflect relative importance, Stallings and Morgan (Biometrika, 2015, in press) introduced general weighted optimality criteria, which assign weights to variances so that greater weight implies greater interest. These criteria are natural extensions of standard design criteria so that design search algorithms can be easily modified to perform optimization with respect to this new class of criteria. This talk first reviews the theory of general weighted optimality criteria and shows how the weighted analogues of standard criteria behave. A straightforward modification of typical design search algorithms is then shown to perform weighted optimization. The algorithm is implemented in SAS PROC OPTEX to find efficient blocked treatment-versus-control designs; unblocked and blocked factorial experiments that are focused on main effect estimation; and factorial experiments under a baseline parameterization.
Oct. 21, 2015
Hassan Allouba :
4 p.m. in SEO 636
Abstract
High order and fractional PDEs have become prominent in theory and in modeling many phenomena. We introduce two large classes of time-fractional and fourth order L-Kuramoto-Sivashinsky (L-KS) Stochastic PDEs. The L-KS PDE/SPDE class is connected to many pattern-formation phenomena. The latter class of time-fractional stochastic equations is related to noisy slow diffusion or diffusion in material with memory. We give comprehensive, sharp, and dimension-dependent Holder and modulus of continuity regularity results for both classes. One important theme of this talk—which is based on a series of our papers—is on the key role Brownian-time processes, their extensions, and their associated kernels play in giving a unifying explicit formulation and in capturing precise behaviors of these two important classes.
Oct. 28, 2015
Yuan Ji :
4 p.m. in SEO 636
Abstract
The Cancer Genomes Atlas (TCGA) data are unique in that multimodal measurements across genomics features, such as copy number, DNA methylation, and gene expression, are obtained on matched tumor samples. The multimodality provides an unprecedented opportunity to investigate the interplay of these features. Graphical models are powerful tools for this task that address the interaction of any two features in the presence of others, while traditional correlation- or regression-based models cannot. We introduce Zodiac, an online resource consisting of a large database containing nearly 200 million interaction networks of multiple genomics features produced by applying novel Bayesian graphical models on TCGA data through massively parallel computation. Setting a new way of integrating TCGA data, Zodiac, publically available at http://www.compgenome.org/ZODIAC, is expected to facilitate the generation of new knowledge and hypotheses by the community.
Nov. 4, 2015
Melanie Pivarski :
4 p.m. in SEO 636
Abstract
Random walks and diffusions are intimately connected through their relationship to the heat (diffusion) equation. We will
look at their structures as well as results comparing large scale asymptotics for finitely generated groups and associated
manifolds (Pittet & Saloff-Coste 2000) and local Dirichlet spaces (P. 2012).
Nov. 18, 2015
Yi Lin :
4 p.m. in SEO 636
Abstract
Fisher information plays an important role in statistics, from asymptotic efficiency, to optimal design of experiments,
to the construction of default priors for Bayesian analysis. However, existence of Fisher information requires regularity conditions.
What happens if the regularity conditions are not met? Is there an alternative measure of information that can
be used in non-regular problems when the Fisher information does not exist? In this talk, I will present a generalization of
the Fisher information to non-regular problems, based on the Hellinger distance, and discuss its properties and some examples. Hints
about its application to optimal design of experiments may also be given.
Dec. 2, 2015
Xia Chen :
4 p.m. in SEO 636
Abstract
This work is concerned with the precise spatial asymptotic behavior for the parabolic Anderson equation
$$\frac{\partial u}{\partial t}(t,x)=\frac{1}{2}\triangle u(t,x)+V(t,x)u(t,x),\quad\quad\mathrm{with}\ u(0,x)=u_0(x),$$
where the homogeneous generalized Gaussian noise $V(t,x)$ is, among other forms, white or fractional white in time
and space. Associated with the Cole-Hopf solution to the KPZ equation, in particular, the precise asymptotic form
$$\lim_{R\to+\infty}(\log R)^{-2/3}\log\max_{|x|\leq R}\, u(t,x)=\frac{3}{4}\sqrt[3]{\frac{2t}{3}}\quad \mathrm{a.s.}$$
is obtained for the parabolic Anderson model $\partial_t u=\frac{1}{2}\partial^2_{xx}u+\dot{W}u$ with the $(1+1)$-white noise $\dot{W}(t,x)$.
Jan. 20, 2016
Jie Yang :
4 p.m. in SEO 636
Abstract
The explosion in the availability of biomedical data is
creating both great opportunities and challenges for collaborative
research among clinicians, genomics and proteomics scientists,
molecular biologists, and statisticians. On one hand, electronic
medical records and genomic data of a large cohort of individuals
are assembled and become available for health study researches. On
the other hand, the combined data are extremely high-dimensional
and also becoming bigger and bigger, especially the genomic part.
As one of the most critical application areas with the biomedical
big data, precision medicine refers to precisely classifying
individuals into subpopulations according to their susceptibility
to a particular disease and precisely tailoring of medical
treatments to subcategories of the disease. Achieving the goals
of precision medicine requires combining data across multiple
formats and developing novel, sophisticated statistical methods.
Our permanental classification approach recently developed
is capable of handling high-dimensional classification problems.
It provides a promising solution for biomedical high-dimensional
data, implemented using the most popular open source statistical
software, R.
Jan. 27, 2016
Ryan Martin :
4 p.m. in SEO 636
Abstract
A Bayesian approach provides a technically straightforward procedure to produce inference on high- and even
infinite-dimensional parameters in complex models. Of course, the choice of a prior is always an issue and,
especially in high-dimensional problems, the prior has a non-trivial effect. One attempt use data to help select
an appropriate prior is <i>empirical Bayes</i> but, unfortunately, this approach does not lead to any
theoretical guarantees that the posterior will behave properly. In this talk I will introduce a very simple strategy
that incorporates data into the prior in such a way that the corresponding posterior distribution has optimal, even
adaptive, concentration rates. Some illustrations of the general theory will also be presented. (This is joint
work with Stephen Walker at University of Texas--Austin.)
Feb. 3, 2016
Xiaoqin Guo :
4 p.m. in SEO 636
Abstract
The Einstein relation describes the relation between the response of a system to a perturbation and its diffusivity at equilibrium. It states that the derivative (with respect to the strength of the perturbation) of the velocity equals the diffusivity. In this talk we consider random walks in iid random conductances on the integer lattice $Z^d$. We show that when $d\ge 3$, the invariant measure for the environment viewed from the particle has a first order expansion in terms of the perturbation. The Einstein relation will follow as a corollary of this expansion. This talk is based on a joint work with N. Gantert and J. Nagel.
Feb. 17, 2016
Igor Cialenco :
4 p.m. in SEO 636
Abstract
We consider a parameter estimation problem for finding the drift coefficient for a large class of parabolic Stochastic PDEs driven by additive or multiplicative noise. In the first part of the talk, we derive several different classes of estimators based on the first N Fourier modes of a sample path observed continuously on a finite time interval. In the second part of the talk we will investigate the simple hypothesis testing problem for the drift coefficient for stochastic fractional heat equation driven by additive noise. We introduce the notion of asymptotically the most powerful test, and find explicit forms of such tests in two asymptotic regimes: large time asymptotics, and increasing number of Fourier modes. Also, we will discuss how to estimate and control the Type~I and Type~II errors. Finally, we illustrate the theoretical results by some numerical examples/simulations.
Feb. 24, 2016
Troy Hernandez :
4 p.m. in SEO 636
Abstract
There's been a lot of hand-wringing over the fumbling of the statistics field in harnessing the
popularity of data science and the "big data revolution". This was most famously addressed in the
editorial from the previous ASA president Marie Davidian, ["Aren't We Data
Science?"](http://magazine.amstat.org/blog/2013/07/01/datascience/). While I have to continue to inform
peers and colleagues that R is indeed a Turing complete programming language (and not just a statistical
computing environment), the data science hype has created technologies that enable any statistician get
into the data science game; most notably Shiny, and cloud-based databases. My talk will focus on weaving
these tools together in a way that allows statisticians to say, ["All your data science are belong to
us."](http://knowyourmeme.com/memes/all-your-base-are-belong-to-us) Additionally, I will provide an
update on the issue of diversity in the tech industry from the frontline.
March 2, 2016
Donald Hedeker :
4 p.m. in SEO 636
Abstract
Intensive longitudinal data are increasingly encountered in many research areas. For example, ecological momentary assessment and/or experience sampling methods are often used to study subjective experiences within changing environmental contexts. In these studies, up to 30 or 40 observations are usually obtained for each subject over a period of a week or so. Because there are so many measurements per subject, one can characterize a subject's mean and variance and can specify models for both. In this presentation, we focus on an adolescent smoking study using ecological momentary assessment where interest is on characterizing changes in mood variation. We describe how covariates can influence the mood variances and also extend the statistical model by adding a subject-level random effect to the within-subject variance specification. This permits subjects to have influence on the mean, or location, and variability, or (square of the) scale, of their mood responses. These mixed-effects location scale models have useful applications in many research areas where interest centers on the joint modeling of the mean and variance structure.
March 9, 2016
Samy Tindel :
4 p.m. in SEO 636
Abstract
We are interested in this talk in ordinary differential equations with a noisy term and a diffusion type coefficient of the form $|x|^{a}$, with a constant a smaller than 1.
This kind of equation has a long story in stochastic analysis. We will first review some of the efforts made by Yamada and Watanabe in this direction (when the equation is driven by a Brownian motion), as well as more some recent developments concerning stochastic PDEs.
We will then introduce two extensions of Young's integral which allows to handle the case of equations driven by a Gaussian signal whose paths are Hölder continuous with Hölder exponent greater than 1/2. Notice that only existence results are obtained, the uniqueness part being still widely open.
This presentation is based on a joint work with J. León and D. Nualart.
March 16, 2016
Rich Bu :
4 p.m. in SEO 636
Abstract
There are two topics in this talk.
1. Applications of statistical models in the consumer lending industry
2. The job market in a nutshell in consumer lending modeling
March 30, 2016
Jyotirmoy Sarkar :
4 p.m. in SEO 636
Abstract
We consider a symmetric random walk on the vertices of a tetrahedron or an octahedron.
Starting from the origin, at each step the random walk moves to one of the vertices adjacent
to the current vertex with equal probability. We find the distribution, or at least the mean
and the variance, of the number of steps needed to (1) return to origin, (2) visit all vertices,
and (3) return to origin after visiting all vertices. We also obtain the distributions of
(i) the number of vertices visited before return to origin, (ii) the last vertex visited, and
(iii) the number of vertices visited while returning to origin after visiting all vertices.
April 6, 2016
Hsin-Hsiung Huang :
4 p.m. in SEO 636
Abstract
We develop a clustering algorithm which does not requires knowing the number
of clusters in advance. Furthermore, our clustering method is rotation-, scale- and
translation-invariant. We call it ``Affine-invariant Bayesian (AIB) process".
A highly efficient split-merge Gibbs sampling algorithm is proposed. Using the
Ewens sampling distribution as prior of the partition and the profile residual
likelihoods of the responses under three different covariance matrix structures, we
obtain inferences in the form of a posterior distribution on partitions.
The proposed split-merge MCMC algorithm successfully and efficiently
estimate the
partition. Our experimental results indicate that the AIB process outperforms
other competing methods. In addition, the proposed algorithm is irreducible
and
aperiodic, so that the estimate is guaranteed to converge to the true
partition.
April 20, 2016
Liang Hong :
4 p.m. in SEO 636
Abstract
Accurate prediction of future claims is a fundamentally important problem in insurance. The Bayesian approach is natural in this context, as it provides a complete predictive distribution for future claims. The classical credibility theory provides a simple approximation to the mean of that predictive distribution as a point-predictor, but this approach ignores other features of the predictive distribution, such as spread, that would be useful for decision-making. Unfortunately, these other features are more sensitive to the choice of loss model and prior distribution, so a flexible nonparametric Bayesian model is desirable. In this paper, we propose a Dirichlet process mixture of log-normals model and discuss the theoretical properties and computation of the corresponding predictive distribution. Numerical examples demonstrate the benefit of our model compared to some existing insurance loss models, and an R code implementation of the proposed method is also provided.
April 27, 2016
Stat faculty and students :
4 p.m. in SEO 636
Abstract
This panel discussion will give students an opportunity to have questions about various things (e.g., advantages and disadvantages of an academic career, general strategies for successful research, career opportunities outside of academia and how to prepare for them, etc) answered by faculty and senior graduate students.
Aug. 31, 2016
Stat faculty and graduate students :
4 p.m. in SEO 636
Sept. 7, 2016
Andrey Sarantsev :
4 p.m. in SEO 636
Abstract
Consider a finite or infinite system of Brownian particles on the real line. Each particle moves as a Brownian motion
with drift and diffusion coefficients depending on its current rank relative to other particles. These systems were
introduced in Banner, Fernholz, Karatzas (2005). Since then, extensive theory was developed for finite systems.
However, infinite systems proved to be much more difficult. We survey the latest results.
Sept. 14, 2016
Bikas Sinha :
4 p.m. in SEO 636
Abstract
Abstract : With reference to 'point estimation' of a real-valued parameter $\theta$ involved in the distribution of a
real-valued random variable $X$, we consider a sample size $n$ and an underlying exact-sense unbiased
estimator ${\hat{\theta}}_n$ of $\theta$ for every $n = k, k+1, k+2, ...\ldots$ where $k$ is the minimum sample
size for existence of an exact-sense unbiased estimator of $\theta$. We wish to investigate exact small sample
properties of the sequence of estimators considered here. This we study by considering what is termed as
'Coverage Probability (CP)' and defined as $CP(n, c)=P[-c < {\hat{\theta}}_n - \theta < c]$. It is desired that the
sequence $[CP(n, c); n=k, k+1, k+2, ...\ldots]$ behaves like an increasing sequence for every $c>0$. We may
note that we are asking for a property beyond 'consistency' of a sequence of estimators. In this presentation we
will discuss several interesting features of the behavior of the $CP(n, c)$.
Sept. 21, 2016
Yan Chang :
4 p.m. in SEO 636
Abstract
I will go over some of principles and typical approaches of credit risk management in a financial company.
Sept. 28, 2016
Frederick Phoa :
4 p.m. in SEO 636
Abstract
Nature-inspired metaheuristic methods, like the particle swarm optimization and many others, enjoys fast convergence towards optimal solution via a series of inter- particle communication. Such methods are common for the optimization problem in engineering, but few in statistics problem. It is especially difficult to implement in some fields of statistics as the search spaces are mostly discrete, while most natural heuristic methods require continuous search domains. This talk introduces a new method called the Swarm Intelligence Based (SIB) method for optimization in statistics problems, featuring the searches within discrete space. Such fields include experimental designs, community detection, change-point analysis, variable selection, etc. The SIB method is a nature-inspired metaheuristic method that includes several operations. This method is advantageous over the traditional particle swarm optimization and many other heuristic approaches in the sense that it is ready for the search of both continuous and discrete domains, and its global best particle is guaranteed to monotonically move towards the optimum. The SIB method is demonstrated in several examples. Several extensions from the standard framework are also discussed at the end of this talk.
Oct. 5, 2016
George Karabatsos :
4 p.m. in SEO 636
Abstract
Dirichlet process (DP) mixture models, as well as models
with mixture distribution assigned a general Bayesian nonparametric (BNP)
prior distribution on the space of probability measures, are
widely-applied and flexible models that can provide reliable statistical
inferences complex data. For such Bayesian mixture models, in practice,
posterior inferences are usually conducted using MCMC, which however, is
prohibitively slow for large data sets. Also for such models, prior
specification can be non-trivial in practice. As alternatives to MCMC, I
consider two new approaches to fast and approximate BNP inference for
large data sets. First, I show that if the ordinary least-squares (OLS)
estimator of the linear regression coefficients is specified as a
functional of the DP posterior distribution, then this functional has
posterior mean given by an observation-weighted ridge regression
estimator, with ridge (coefficient shrinkage) parameter given by the DP
precision parameter;
and has a heteroscedastic-consistent posterior covariance matrix.
This result is based on the multivariate delta method applied to
prior-informed bootstrap distribution approximation to the DP posterior.
Second, I consider an approximation to the BNP (infinite) mixture model
that I introduced and studied in several articles, defined by ordinal
regression mixture weights.The approximate model is defined by a (large)
finite mixture, with each component distribution multiplied by a histogram
bin indicator function. I show that posterior inference with this
approximate BNP model can be conducted by iteratively-reweighted least
squares estimation for the mixture weight parameters, and least-squares
estimation for the component densities, all involving computations that
are orders of magnitude faster that MCMC-based inference of the original
mixture model. This is also true for a version of the approximate model
that is defined by an ordinal regression of DPs. I illustrate the two
approximate BNP methods through the analysis of real data sets.
Oct. 12, 2016
Tai Melcher :
4 p.m. in SEO 636
Abstract
As in the Riemannian setting, a subRiemannian heat kernel is controlled by the geometry of the underlying manifold. In particular, the asymptotic behavior of the kernel can reveal certain geometric and topological data. We study the logarithmic derivatives of subRiemannian heat kernels in some cases and show that, under appropriate scaling, they converge to their analogues on stratified groups. This gives one quantification of the now standard idea that stratified groups play the role of the tangent space to subRiemannian manifolds.
This is joint work with Joshua Campbell.
Oct. 19, 2016
Dongxiao Zhu :
3 p.m. in SEO 636
Abstract
In multi-class classification, different classes may relate to different feature groups. In this
talk, I will present a class-conditional regularization of the multinomial logistic model to enable the
discovery of class-specific feature groups. I will also present an efficient cyclic block coordinate descent
based algorithm to solve the model. In another work, I will introduce a novel joint mixture model framework
to estimate cluster size distribution, particularly for over-dispersed (high variance) ones, together with
cluster compactness (density). Our methods are sufficiently flexible and general to be applied to multiple
application domains, such as social networks, image segmentation, natural language processing and
bioinformatics.
Milan Stehlik :
4 p.m. in SEO 636
Abstract
Since 2004 there were many results obtained regarding the determination of optimal designs for models with
correlated errors. This task is substantially more difficult than in case of iid errors and for this reason
not so well developed. Stochastic process with parametrized mean and covariance is observed over a compact
set. The information obtained from observations is measured through the information functional (defined on
the Fisher information matrix). The role of equidistant designs has been recognized; e.g. such designs have
been proved to be optimal for parameter of trend of stationary Ornstein-Uhlenbeck process, also for
nonstationary Ornstein-Uhlenbeck process, both for prediction and estimation. We can conclude that if
only trend parameters are of interest, the designs covering more-less uniformly the whole design space
are rather efficient when correlation decreases exponentially. This concept is also valid for so called
monotonic set designs. We will concentrate on several important issues regarding regularity conditions for
quality of ``plug-in" approach from iid case. Namely, 1) relaxing the continuity of covariance. We will
introduce the regularity conditions for isotropic processes with semicontinuous covariance such that
increasing domain asymptotic is still feasible, however more flexible behavior may occur here.
In particular, the role of the nugget effect will be illustrated. 2) regarding quality of
approximation of inverse information matrix by variance-covariance, Pazman (2007) discusses
theoretical background and formulated important conditions, which are fundamental.
Zhu and Stein (2005) made simulations experiments. Finally, application in troposphere
methane modelling will be illustrating the developed methods.
Oct. 26, 2016
Lihui Zhao :
4 p.m. in SEO 636
Abstract
For a study with an event time as the endpoint, its survival function contains all the information
regarding the temporal, stochastic profile of this outcome variable. The survival probability at a
specific time point, say t, however, does not transparently capture the temporal profile of this
endpoint up to t. An alternative is to use the restricted mean survival time (RMST) at time t to
summarize the profile. The RMST is the mean survival time of all subjects in the study population
followed up to t, and is simply the area under the survival curve up to t. The advantages of using
such a quantification over the survival rate have been discussed in the setting of a fixed-time
analysis. In this research, we generalize this approach by considering a curve based on the RMST over
time as an alternative summary to the survival function. Inference, for instance, based on
simultaneous confidence bands for a single RMST curve and also the difference between two RMST curves
are proposed. The latter is informative for evaluating two groups under an equivalence or
noninferiority setting, and quantifies the difference of two groups in a time scale. In addition, we
extend RMET to the setting of multiple endpoints, which includes classical competing risks and
semi-competing risks. The methods are illustrated with the data from two clinical trials.
Nov. 2, 2016
Mike Cranston :
4 p.m. in SEO 636
Abstract
We consider large time behavior of typical paths under the Anderson polymer measure. If $P^x_\kappa$ is the measure induced by rate $\kappa,$ simple, symmetric random walk on $\mathbb{Z}^d$ started at $x,$ this measure is defined as
\[d\mu^x_{\kappa,\beta,T}T(X)={Z_{\kappa,\beta,T}}^{-1} \exp\left\{\beta\int_0^T dW_{X(s)}(s)\right\}dP^x_\kappa(X)\]
where $\{W_x:x\in \mathbb{Z}^d\}$ is a field of $iid$ standard, one-dimensional Brownian motions, $\beta>0, \kappa>0$ and
$Z_{\kappa,\beta,t}(x)$ the normalizing constant.
We establish that the polymer measure gives a macroscopic mass to a small neighborhood of a typical path as $T \to \infty$, for parameter values outside the perturbative regime of the random walk, giving a pathwise
approach to polymer localization, in contrast with existing results. The localization becomes complete as $\frac{\beta^2}{\kappa}\to\infty$ in the sense that the mass grows to 1.
The proof makes use of the overlap between two independent samples drawn under the Gibbs measure $\mu^x_{\kappa,\beta,T}$,
which can be estimated by the integration by parts formula for the Gaussian environment.
Conditioning this measure on the number of jumps, we obtain a canonical measure which already shows scaling
properties, thermodynamic limits, and decoupling of the parameters. This talk is based on joint work with Francis Comets.
Nov. 9, 2016
Sanjib Basu :
4 p.m. in SEO 636
Abstract
We consider the question of variable selection in complex models. This is often a difficult problem due to the inherent nonlinearity of the models and the resulting non-conjugacy in their Bayesian analysis. Bayesian variable selection in lifetime data models often utilize cross-validated predictive model selection criteria which can be relatively easy to estimate for a given model. However, the performances of these criteria are not well-studied in large-scale variable selection problems and, evaluation of these criteria for each model under consideration can be difficult to infeasible. An alternative criterion is based on the highest posterior model but its implementation is difficult in non-conjugate lifetime models. In this presentation, we compare the performances of these different criteria in complex lifetime data models including models with limited failure. We also propose an efficient variable selection method and illustrate its performance in simulation studies and real example
Nov. 16, 2016
Si Tang :
4 p.m. in SEO 636
Abstract
The classic Susceptible-Infected-Resistant (SIR) model, due to Kermack & McKendrick (1927), describes the spread of an infectious disease in an infinite, homogeneous population using a system of ordinary differential equations. In this talk, I will focus on stochastic SIR models, where the population is of size $N$ and transmissions only occurs locally. In these models, the sizes of infected clusters are characterized by the excursion lengths of a continuous stochastic process, denoted by $W_t$. In particular, in the mean-field SIR case, $W_t$ is a reflected (at 0) Brownian motion with negative drift; in the spatial case, $W_t$ is a reflected Brownian motion trimmed by a Poisson point process whose intensity is determined by a super Brownian motion with time-and-location dependent killing.
Nov. 30, 2016
Ella Revzin :
4 p.m. in SEO 1227
Abstract
This talk will introduce students to Precima and to Marketing Data Science. Precima is a Marketing Analytics company that specializes in data driven products and services that help retailers and manufacturers drive sales growth and boost profitability. I will provide an overview of the company and its Data Science practice, with a focus on Targeted Marketing. For the latter, I’ll review a typical data acquisition and modeling process. Finally, I’ll describe a typical week in the life of a Precima Data Scientist and give information on what we look for when we hire for internships and full time positions.
Feb. 15, 2017
Jinqiao Duan :
4 p.m. in SEO 636
Abstract
Dynamical systems arising in engineering and science are often subject to random fluctuations. The noisy fluctuations may be Gaussian or non-Gaussian, which are modeled by Brownian motion or α-stable Levy motion, respectively. Non-Gaussianity of the noise manifests as nonlocality at a “macroscopic” level. Stochastic dynamical systems with non-Gaussian noise (modeled by α-stable Levy motion) have attracted a lot of attention recently. The non-Gaussianity index α is a significant indicator for various dynamical behaviors.
The speaker will overview recent advances in non-Gaussian stochastic dynamical systems, highlighting deterministic and numerical methods, including analysis and simulation of mean exit time and escape probability. Some materials are taken from the speaker's new book “An Introduction to Stochastic Dynamics” (Cambridge University Press, 2015) .
Feb. 22, 2017
Tingting Cheng :
4 p.m. in SEO 636
Abstract
We study a functional coefficient time series model with trending regressors, where the coefficients are unknown functions of time and random variables. We propose a local linear estimation method to estimate the unknown coefficient functions. An asymptotic distribution of the proposed local linear estimator is established under mild conditions. A test procedure is developed to test the null hypothesis that the functional coefficients take particular parametric forms. For practical use, we further propose a Bayesian approach to select bandwidths involved in this local linear estimator. Several numerical examples are provided to examine the finite sample performance of the proposed local linear estimator and the test procedure. The results show that the local linear estimator works well and the proposed test has satisfactory size and power. In addition, simulation studies show that the Bayesian bandwidth selection method is better than cross–validation method. Furthermore, we employ the functional coefficient model to study the relationship between consumption per capita and income per capita in U.S. and the results show that functional coefficient model with our proposed local linear estimator and Bayesian bandwidth selection method performs best in both in–sample fitting and out–of–sample forecasting.
March 29, 2017
Dr. Viswanath Devanarayan :
4 p.m. in SEO 636
Abstract
In High-Throughput-Screening efforts during the drug discovery process, hundreds of thousands of compounds are tested to identify promising drug candidates that modulate specific gene targets. These drug candidates may ultimately serve as therapeutic candidates for some disease indications of interest. Critical decisions related to compound selection and prioritization are made based on fairly limited data, and therefore rely greatly on data quality and reproducibility. Standard statistical metrics and methods in textbooks do not directly apply for these evaluations. This presentation will provide an overview of some statistical measures that were developed specifically for this application. The content of this presentation will be very practical and data-driven, and hence will be suitable for a broad audience.
April 5, 2017
Karl Liechty :
4 p.m. in SEO 636
Abstract
It is a well known and celebrated fact that the eigenvalues of random Hermitian matrices from a unitary invariant ensemble form a determinantal point process with correlation kernel given in terms of a system of orthogonal polynomials on the real line. It is a much more recent result that the eigenvalues of the sum of such a random matrix with a matrix from the Gaussian unitary ensemble (GUE) also forms a determinantal point process, with the kernel given in terms of the Weierstrass transform of the original kernel. I'll talk about the case in which the limiting distribution of eigenvalues is critical in the sense that there is a non-generic scaling limit for the correlation kernel, and discuss the effect of a Gaussian perturbation on the limiting critical kernel. This is joint work with Tom Claeys, Arno Kuijlaars, and Dong Wang.
April 7, 2017
Wei Zheng :
11 a.m. in TH300
April 12, 2017
Ju-Yi Yen :
4 p.m. in SEO 636
Abstract
In this talk, we study the process obtained from a Brownian bridge after excising all the excursions below the waterline level which reach zero. Three variables of interest are the maximum of this process, the value where this maximum is attained, and the total length of the excursions which are excised. Our analysis relies on some interesting transformations connecting Brownian path fragments and the 3-dimensional Bessel process.
April 19, 2017
Xi Geng :
4 p.m. in SEO 636
Abstract
The exponential transform of a vector-valued path, also known as the signature of a path, is the formal sequence of associated iterated path integrals. It is widely believed (and surprisingly) that the signature contains essentially all information about the underlying path. In this talk, we will prove that every (rough) path is uniquely determined by its signature up to certain tree-like equivalence. Moreover, looking into its probabilistic counterpart, we will obtain stronger uniqueness results for sample paths of Gaussian processes by applying the technique of Malliavin's calculus. This part inspires the development of a universal way to reconstruct every rough path from its signature.
April 26, 2017
Prof. Hsin-Hsiung Huang :
4 p.m. in SEO 636
Abstract
Support Vector Machines (SVM) classifier is a popular classification method.
However, most users may not well take tuning parameters selection because
this step is time consuming. In practice, the tuning parameters are chosen
by evaluating parameter candidates via cross validation. It is shown that
the performance of SVM is sensitive to the values of tuning parameters. In
some cases, SVM performs poorly due to the values of tuning parameters.
However, selection of parameter values for SVM often relies on inefficient approaches
such as extensive cross validation. To get around the problem, users
may resort to anecdotal methods or default values set by software developers.
However, these methods may compromise performance of classification accuracy.
In this research, we propose an efficient algorithm called P-SVM for
selecting the parameter pair, (gamma,C), of SVM with Gaussian kernels on metric
data. P-SVM searches only a handful of percentiles of the squared Euclidean
distances of data points to select the best pair of parameter values. Our motivation
case study of business intelligence categorization demonstrates that
P-SVM achieved a signi cant improvement in precision, recall, F-measure,
and AUC from the default parameter values settled in Weka, a widely used
data mining software. Applications of both simulation and publicly-available
datasets also demonstrate that P-SVM achieves substantial improvement in
computational time without loss of much classification accuracy.
May 3, 2017
Prof. Wei Zheng :
4 p.m. in SEO 636
Abstract
Crossover design is a design of experiments, where a subject receives a
sequence of various treatment over a period of time points. While it provides the
within subject comparison between treatment effects, the potential carryover effect
in the model makes the study of optimal crossover designs quite complicated. Such
study was initiated by Hedayat and Afsarinejad (1978), and many researchers have
contributed to the general theory for optimal designs. Among them, Kushner (1997)
developed very elegant results for the optimality conditions in the approximate
design theory. My talk will mainly focus on this approach and talk about some recent
progress as well as future challenges. I will also share some of my own thoughts of
how to tackle these problems.
Aug. 7, 2017
Dr. Lan Xue :
3 p.m. in SEO 636
Abstract
Missing data is one of the major methodological problems in longitudinal studies. It not only reduces the sample size, but also
can result in biased estimation and inference. It is crucial to correctly understand the missing mechanism and appropriately
incorporate it into the estimation and inference procedures. Traditional methods, such as the complete case analysis and imputation
methods, are designed to deal with missing data under unverifiable assumptions of MCAR and MAR. The purpose of this talk is to identify
and estimate missing mechanism parameters under the non-ignorable missing assumption utilizing the refreshment sample. In particular, we
propose a semi-parametric method to estimate the missing mechanism parameters by comparing the marginal density estimator using Hirano?s
two constraints (Hirano et al. 1998) along with additional information from the refreshment sample. Asymptotic properties of semi-parametric estimators are developed. Inference based on bootstrapping is proposed and verified through simulations.
Aug. 30, 2017
Yinghui Shi :
4 p.m. in SEO 636
Abstract
The intrinsic ultracontractivity (Abbr. I.U.)
of the subprocesses $X_D^b$ of two kinds of special Markov processes $X^b$ upon leaving any bounded open set $D \subset \mathbb{R}^d$
will be given in this talk.
Here the processes $X^b$ are associated with the operator $\mathcal{L}^b= \Delta^{\alpha/2} +\mathcal{S}^b$ with
$d \geq 1$ and $0 < \beta < \alpha \leq 2$, where
$$\mathcal{S}^bf(x):=\int_{\mathbb{R}^d}(f(x+z)-f(x)-\nabla f(x)\cdot z \mathbb{1}_{\{|z|\leq 1\}})\frac{b(x,z)}{|z|^{d+\beta}}dz$$
and $b(x,z)$ is a bounded Borel function on $\mathbb{R}^d\times \mathbb{R}^d$ with $b(x,z)=b(x,-z)$ for $x,z\in \mathbb{R}^d$.
The operator $\mathcal{L}^b$ can be seen as the Laplacian ($\alpha = 2$) or the fractional Laplace operator ($0<\alpha <2$)
with a lower order perturbation $\mathcal{S}^b$.
Our main results are proved under the frame of the I.U. for non-symmetric Levy
processes. We discuss the transition density function for $X_D^b$ firstly and then we get its dual process
under some reference measure. At last, we prove that I.U. stands under the conditions as follows: for any compact subset $K, L \subset\mathbb{R}^d$,
$\inf_{x\in K}\inf_{z\in L}b(x,z)>0$ in the case of $\alpha=2$ and $\inf_{x\in K}\inf_{z\in L}(1+\frac{b(x,z)}{\mathcal{A}(d,\alpha)})|z|^{\alpha-\beta}>0$
in the case of $0<\alpha<2$. (Joint work with Yi, Bingji and Song, Renming)
Sept. 6, 2017
Stat faculty and graduate students :
4 p.m. in SEO 636
Abstract
TBA
Sept. 13, 2017
Prof. Zhen Liu :
4 p.m. in SEO 636
Abstract
We study the classic inventory pooling problem by Eppen (1979) under a special
class of multivariate fat-tail distribution: Normal Inverse Gaussian (NIG) demands to
better fit real-world demand data. We obtain the optimal inventory level in a closed
form by employing standardized NIG density function, and express the optimal expected
costs in terms of unit NIG loss function. In addition to independent and identically
distributed demands, our results complement Bimpikis and Markakis (2015) by considering
correlated demands. We further discuss the transshipment problem of Dong and Rudi (2004)
under NIG demands.
Sept. 27, 2017
Yanghui Liu :
4 p.m. in SEO 636
Abstract
The term “limit theorem” is associated with a multitude of statements having to do with the convergence of probability distributions of sums of increasing number of random variables. Given that a limit theorem result holds, “weighted limit theorem” considers the asymptotic behavior of the corresponding weighted sums. The weighted limit theorem problem has drawn a lot of attention in recent articles due to its key role in topics such as parameter estimations, Ito’s formula in law, time-discrete numerical schemes, and normal approximations, and various “unexpected” weighted limit theorems have been discovered since then. The purpose of this talk is to introduce a general framework and a transferring principle for this problem, and to provide improvement of the existing results in a few aspects.
Oct. 4, 2017
Yimin Xiao :
4 p.m. in SEO 636
Abstract
Multivariate (or vector-valued) stochastic processes are important in probability, statistics and various scientific areas as stochastic models. In recent years, there has been increasing interest in investigating their statistical inference and prediction.
In this talk, we study the problem for estimating jointly the fractal indices of a bivariate Gaussian process. These indices not only determine the smoothness of each component process, fractal behavior of the whole process, but also play important roles in characterizing the dependence structure among the components.
Under the infill asymptotics framework, we establish joint asymptotic results for the increment-based estimators for bivariate fractal indices. Our main results show the effect of the cross dependence structure on the performance of the estimators.
This is a joint paper with Yuzhen Zhou.
Oct. 18, 2017
Dan Spillane :
4 p.m. in SEO 636
Abstract
What's the purpose of Data Science anyway? In this discussion we'll explore how we need to turn data science upside-down to create the real value of this powerful trade. We need to push desired (business, social, economic...) outcomes to the forefront (the hypothesis) and leverage data, data platforms and AI to develop the questions we don't even know to ask and then help answer. We need to be data pioneers not just data engineers. Looking forward to a fruitful and living dialogue on Data Science 2.0.
Oct. 25, 2017
Sayar Karmakar :
4 p.m. in SEO 636
Abstract
The term "time-varying(tv) coefficient model" refers to the framework of time series and regression models where the unknown coefficients vary across time. Deviation from the constancy of parameters is more natural due to the effect of several external factors/events or sometimes simply for a very long time-horizon. For the past two decades, tv regression models received considerable attentions however not much was done for conditional heteroscedastic(CH) models until very recently.
The purpose of this talk is to introduce an unanimous framework to combine the treatments for tv linear regression, tv generalized regression and tv time-series models. Local linear M-estimation is used to estimate the unknown curves. We obtain a Bahadur representation of these estimated curves and use it to find the simultaneous confidence bands. To circumvent the logarithmic convergence rate of the theoretical bands, a Bootstrap method is proposed using an optimal Gaussian approximation. Some simulations for tvARCH and tvGARCH models and analysis of some stock market datasets are presented.
This is a joint work with Stefan Richter and Wei Biao Wu.
Nov. 1, 2017
Yunxiao He :
4 p.m. in SEO 636
Abstract
The marketing industry produces and consumes an enormous amount of data. This talk will provide a quick overview of the marketing industry with focus on illustrating how data and statistical tools can be leveraged in connecting brands and consumers.
Nov. 22, 2017
Xi Geng :
3 p.m. in SEO 636
Abstract
In the groundbreaking work of B. Hambly and T. Lyons (Uniqueness for the signature of a path of bounded variation and the reduced path group, Ann. of Math., 2010), it has been conjectured that the geometry of a tree-reduced bounded variation path can be recovered from the tail asymptotics of its associated sequence of iterated path integrals. While this conjecture is still remaining open in the general deterministic case, in this talk we investigate a similar problem in the probabilistic setting for Brownian motion. It turns out that a martingale approach applied to the hyperbolic development of Brownian motion allows us to extract useful information from the tail asymptotics of Brownian iterated integrals, which can be used to determined the Brownian rough path along with its natural parametrization uniquely. This in particular strengthens the existing uniqueness results in the literature.
Lei Liu :
4 p.m. in SEO 636
Abstract
In many biomedical studies, disease progress is monitored by a biomarker over time, e.g., repeated measures of CD4, hemoglobin level in end stage renal disease (ESRD) patients. The endpoint of interest, e.g., death or diagnosis of a specific disease, is correlated with the longitudinal biomarker. The causal relation between the longitudinal and time to event data is of interest. In this paper we examine the causality in the analysis of longitudinal and survival data. We consider four questions: (1) whether the longitudinal biomarker is a mediator between treatment and survival outcome; (2) whether the biomarker is a surrogate marker; (3) whether the relation between biomarker and survival outcome is purely due to an unknown confounder; (4) whether there is a mediator moderator for treatment. We illustrate our methods by data from two clinical trials: an AIDS study and a liver cirrhosis study.
Dec. 6, 2017
Hongmei Jiang :
4:15 p.m. in SEO 636
Abstract
Metagenomics is a powerful tool to study the microbial organisms living in various environments. The abundance of a microorganism or a taxon is usually estimated using relative proportion or percentage in sequencing-based metagenomics studies. Due to the constraint of the sum of the relative abundances being 1 or 100%, standard conventional statistical methods may not be suitable for metagenomics data analysis. In this talk we will discuss characterization of the association between microbiome and disease status and variable selection in regression analysis with compositional covariates. Current statistical and computational methods that are being developed to analyze the metagnoimcs data and the challenges will also be highlighted.
Feb. 7, 2018
Lulu Kang :
4 p.m. in SEO 636
Abstract
In many science and engineering systems both quantitative and qualitative output observations are collected. For short, we call such a system QQ system. In this talk, I will talk about a systematical approach for the experimental design and data analysis for the QQ system.
Classic experimental design methods are not suitable here because they often focus on one type of responses. We develop both Bayesian D and A-optimal design methods for experiments with one continuous and one binary responses. Both noninformative and conjugate informative prior distributions on the unknown parameters are considered. The proposed design criterions has meaningful interpretations in terms of the optimality for the models for both types of responses. Efficient design construction algorithms are developed to construct the local D-and A-optimal designs for given parameter values.
To capture a correlation between the two types of responses, we propose a Bayesian hierarchical modeling framework to jointly model a continuous and a binary response. Compared with the existing methods, the Bayesian method overcomes two restrictions. First, it solves the problem in which the model size (specifically, the number of parameters to be estimated) exceeds the number of observations for the continuous response. Second, the Bayesian model can provide statistical inference on the estimated parameters and predictions. Gibbs sampling scheme is used to generate accurate estimation and prediction for the Bayesian hierarchical model. Both simulation and real case study are shown to illustrate the proposed method.
Feb. 28, 2018
Li Wang :
4 p.m. in SEO 636
Abstract
In dose ranging clinical trials, it is critical to investigate the dose-response profile and to identify a minimum effective dose (MED) to guide the dose selection for phase 3 confirmatory trials. Traditional dose ranging trials focus on pairwise comparisons between placebo and each investigational dose, while in recent years MCP-Mod (Multiple Comparison Procedures & Modeling) arose and gained popularity in the design and analysis of dose ranging trials. Comprehensive comparison between MCP-Mod and other methods have been made on continuous variables assuming a normal distribution. We extend the comparison to binary/binomial response variables. Via simulation, the rate of correct and incorrect MED identification are compared for Dunnett's test, trend test and MCP-Mod for a variety of underlying dose response profiles including both monotone and non-monotone dose responses and are compared under a large number of trial design settings. The precision of MED estimation using MCP-Mod is also evaluated comparing the design options of more dose levels and smaller sample size per dose versus fewer dose levels and larger sample size per dose
March 7, 2018
Ruijun Zhao :
4 p.m. in SEO 636
Abstract
Molecular dynamics (MD) is a computer simulation method for studying the physical movement of atoms and molecules.
It has broad applications in many fields of sciences.
In this talk, I will discuss how we use MD to compute transition pathways of conformational change,
given two different metastable states of a biomolecule.
In particular, we proposed an efficient algorithm, Maximum Flux Transition Paths, to compute such a path and applied the method to
the Src tyrosine kinase family, which has long been implicated in the development of cancer.
Maximum Flux Transition Paths relies on efficiently computing the free energy and
proto-diffusion tensor, which are computed as the conditional expectations of
some ``observable" A(x) that depend on random states x drawn from distributions that are known except for their normalizing factor.
Vast amounts of computer time are used to compute these expectations.
Markov chain Monte Carlo methods are very popular for computing these expectations.
In this talk, I will also discuss the challenge of this method and how to estimate the accuracy of these expectations.
March 14, 2018
Yiou Li :
4 p.m. in SEO 636
Abstract
The generalized linear model plays an important role in statistical analysis and the related design issues are undoubtedly challenging. The state-of-the-art works mostly apply to design criteria on the estimates of regression coefficients. It is of importance to study optimal designs for generalized linear models, especially on the prediction aspects.
In this talk, I will discuss a prediction-oriented design criterion, I-optimality, and how we develop an efficient sequential algorithm of constructing I-optimal designs for generalized linear models. Through establishing the General Equivalence Theorem of the I-optimality for generalized linear models, an insightful understanding is obtained for the proposed algorithm on
how to sequentially choose the support points and update the weights of support points of the design. The proposed algorithm is computationally efficient with guaranteed convergence property. Numerical examples are conducted to evaluate the feasibility and computational efficiency of the proposed algorithm.
March 21, 2018
Yuguo Chen :
4 p.m. in SEO 636
Abstract
Random graphs with given vertex degrees have been widely used as a model for many real-world complex networks. We describe a sequential sampling method for sampling networks with a given degree sequence. These samples can be used to approximate closely the null distributions of a number of test statistics involved in such networks, and provide an accurate estimate of the total number of networks with given vertex degrees. We apply our method to a range of examples to demonstrate its efficiency in real problems.
April 4, 2018
Daniel W. Apley :
4 p.m. in SEO 636
Abstract
For many supervised learning applications, understanding and visualizing the effects of the predictor variables on the predicted response is of paramount importance. A shortcoming
of black box supervised learning models (e.g., complex trees, neural networks, boosted trees, random forests, nearest neighbors, local kernel-weighted methods, support vector regression,
etc.) in this regard is their lack of interpretability or transparency. Partial dependence (PD) plots, which are the most popular general approach for visualizing the effects of the predictors with
black box supervised learning models, can produce erroneous results if the predictors are strongly correlated, because they require extrapolation of the response at predictor values that are far outside the multivariate envelope of the training data. Functional ANOVA for correlated inputs can avoid this extrapolation but involves prohibitive computational expense and subjective choice of additive surrogate model to fit to the supervised learning model. We present a new visualization approach that we term accumulated local effects (ALE) plots, which have a
number of advantages over existing methods. First, ALE plots do not require unreliable extrapolation with correlated predictors. Second, they are orders of magnitude less computationally expensive than PD plots, and many orders of magnitude less expensive than functional ANOVA. Third, they yield convenient variable importance/sensitivity measures that
possess a number of desirable properties for quantifying the impact of each predictor.
April 11, 2018
Shunpu Zhang :
4 p.m. in SEO 636
Abstract
We propose a new multiple test called the minPOP test and two of its modified versions (the left truncated and the double truncated minPOP tests) for testing multiple hypotheses simultaneously. We show that these tests have multiple testing procedures based on these tests have strong control of the family-wise error rate. A method for finding the p-values of the proposed multiple testing procedures after adjusting for multiplicity is also developed. Simulation results show that the minPOP tests in general have higher global power than the existing well known multiple tests, especially when the number of hypotheses being compared is relatively large. Among the multiple testing procedures we developed, we find that the ones based on the left truncated and double truncated minPOP tests tend to have higher number of rejections than the existing multiple testing procedures. In the case of correlated test statistics, simulation results show that only the double truncated minPOP test is reasonably robust to positively correlated test statistics, while all the other tests seem to be robust to negatively correlated test statistics.
April 18, 2018
Jean-Pierre Fouque :
3 p.m. in SEO 636
Abstract
Rough stochastic volatility models have attracted a lot of attention recently, in particular for the linear option pricing problem. In this talk, starting with power utilities, we propose to use a martingale distortion representation of the optimal value function for the nonlinear asset allocation problem in a (non-Markovian) fractional stochastic environment (for all Hurst index $H \in (0, 1)$). We rigorously establish a first order approximation of the optimal value, when the return and volatility of the underlying asset are functions of a stationary slowly varying fractional Ornstein-Uhlenbeck process. We prove that this approximation can be also generated by the zeroth order trading strategy providing an explicit strategy which is asymptotically optimal in all admissible controls. Furthermore, we extend the discussion to general utility functions, and obtain the asymptotic optimality of this strategy in a specific family of admissible strategies. If time permits, we will also discuss the problem under fast mean-reverting fractional stochastic environment.
Joint work with Ruimeng Hu (UCSB).
Ruoqing Zhu :
4 p.m. in SEO 636
Abstract
We propose a class of dimension reduction methods for right censored survival data using a counting process representation of the failure process. Semiparametric estimating equations are constructed to estimate the dimension reduction subspace for the failure time model. The proposed method addresses two fundamental limitations of existing approaches. First, using the counting process formulation, it does not require any estimation of the censoring distribution to compensate the bias in estimating the dimension reduction subspace. Second, the nonparametric part in the estimating equations is adaptive to the structural dimension, hence the approach circumvents the curse of dimensionality. Asymptotic normality is established for the obtained estimators. We further propose a computationally efficient approach that requires only a singular value decomposition to estimate the dimension reduction subspace. Numerical studies suggest that the proposed methods exhibit significantly improved performance for estimating the true dimension reduction subspace. We further conducted a real data analysis on a skin cutaneous melanoma dataset from The Cancer Genome Atlas. The findings have important biological implications. The proposed methods are implemented in the R package ``orthoDr'', which efficiently solves the semiparametric estimating equations within the Stiefel manifold of the parameter space.
April 25, 2018
Rui Song :
4 p.m. in SEO 636
Abstract
In the first part of the talk, we propose a new concordance-assisted
learning for estimating optimal individualized treatment regimes. We
first introduce a type of concordance function for prescribing
treatment and propose a robust rank regression method for estimating
the concordance function. We then find treatment regimes, up to a
threshold, to maximize the concordance function, named prescriptive
index. Finally, within the class of treatment regimes that maximize
the concordance function, we find the optimal threshold to maximize
the value function. Although this method makes better use of the
available information through pairwise comparison, the objective
function is discontinuous and computationally hard to optimize. In the
second part of the talk, we consider a convex surrogate loss function
to solve this problem. In addition, our algorithm ensures sparsity of
decision rule and makes it easy to interpret. Simulation results of
various settings and application to STAR*D both illustrate that the
proposed method can still estimate optimal treatment regime
successfully when the numb of covariates is large.
May 2, 2018
Peter Bonate :
3 p.m. in SEO 636
Abstract
Dr. Peter Bonate has over 20 years experience in modeling and simulation in the pharmaceutical industry. Dr. Bonate
will discuss his career and the role modeling and simulation has played in the development of many different pharmaceutical
products.
Aug. 29, 2018
Donald E.K. Martin :
4 p.m. in 636 SEO
Abstract
Higher-order Markov models provide a good approximation to probabilities associated with many categorical time series, and thus they are applied extensively. However, a major drawback associated with them is that the number of model parameters grows exponentially in the order of the model, and thus only very low-order models are considered in applications. Another drawback is lack of flexibility, in that higher-order Markov models give relatively few choices for the number of model parameters. Sparse Markov models are Markov models where transition probabilities are lumped into classes comprised of invariant probabilities. The contexts for conditioning may be either hierarchical (as in variable length Markov chains) or non-hierarchical. This supplies a model that helps with the two problems given above, and which thus gives a better handling of the trade-off between bias associated with having too few model parameters and variance associated with having too many. In this work, methods for efficient computation of pattern distributions through Markov chains with minimal state spaces are extended to the sparse Markov framework.
Sept. 5, 2018
Wei Sun :
4 p.m. in 636 SEO
Abstract
Tensor as a multi-dimensional generalization of matrix has received increasing attention due to its success in many empirical tasks. In particular, dynamic tensor data are becoming prevalent since time is often one of the tensor modes. Existing tensor clustering methods either fail to account for the dynamic nature of the data, or are inapplicable to a general-order tensor. Also there is often a gap between statistical guarantee and computational efficiency for existing tensor clustering solutions. In this talk, I will introduce a new dynamic tensor clustering method, which takes into account both sparsity and fusion structures, and enjoys strong statistical guarantees as well as high computational efficiency. The efficacy of our approach will be illustrated via two real applications: brain dynamic functional connectivity analysis, and online advertisement clustering for market segmentation.
Sept. 12, 2018
Organizational meeting :
4 p.m. in 636 SEO
Sept. 19, 2018
Shuwen Lou :
4 p.m. in 636 SEO
Abstract
We now live in a world surrounded by data. As an example, when we want to buy or sell a house, we browse real estate websites and go through related listings. By comparing the "data", we subconsciously ``generate a price quote" for the house we are interested in buying or trying to sell. This can be viewed as an optimization problem which, in theory, can be solved using gradient descent (GD) method. However, in real-world scenarios, because of the tremendous sizes of the datasets, vanilla GD is typically not an efficient or computable option. An improved version of vanilla GD is stochastic gradient descent (SGD).
In the first half of this talk, we will go through the background of GD and SGD algorithms. From there, we will introduce how a discrete-time SGD algorithm can be modeled by a continuous-time stochastic process based on Brownian motion. Using probabilistic tools, one can reveal many interesting properties from continuous-time versions of SGD. Some of these properties can be translated back to discrete-time SGD algorithms.
Sept. 26, 2018
Kevin Potcner :
4 p.m. in 636 SEO
Abstract
As the size and sources of data becomes more available in today's business environments, data analysts are beginning to add more sophisticated predictive statistical modeling techniques to their analysis toolkit.
A typical real-world predictive modeling workflow includes data cleaning and exploration, model fitting, model validation, model comparison, final model selection and deployment of the final predictive model.
In this presentation, a statistical scientist from JMP will illustrate the predictive modeling workflow by analyzing a real dataset. After data preparation and initial exploration, we will create a number of predictive models such as Multiple Linear Regression, Regression tree, Neural Net, and K-Nearest Neighbors.
We will evaluate each model and select the best model using the Prediction Profiler and JMP's Model Comparison tool.
Code will be automatically created in a variety of programming languages (e.g., SAS, SQL, Python, et al.) in order to implement that model in a production environment.
Oct. 3, 2018
Guanhua Chen :
4 p.m. in 636 SEO
Abstract
We propose a new method termed stabilized O-learning for deriving stabilized dynamic treatment regimes (DTRs), which are sequential decision rules for individual patients not only adapt over the course of the disease progression but also consistent over time in its format. The method provides a robust and efficient learning framework for constructing DTRs by directly optimizing a doubly robust estimator of the expected long-term outcome. It can accommodate various types of outcomes, including continuous, categorical and potentially censored survival outcomes. In addition, the method is flexible to incorporate clinical preferences into a qualitatively fixed rule, where the parameters indexing the decision rules that are shared across stages can be estimated simultaneously. We conducted extensive simulation studies, showing a superior performance of the proposed method. We analyzed the data from the prospective Canary Prostate Cancer Active Surveillance study using the proposed method.
Oct. 10, 2018
Xinyi Li :
4 p.m. in 636 SEO
Abstract
We consider loop-erased random walk (LERW) in three dimensions and give an asymptotic estimate on the one-point function for LERW and the non-
intersection probability of LERW and simple random walk in three dimensions. Then we show that 3D LERW converges to its scaling limit in natural parametrization. This is a joint work in progress with Daisuke Shiraishi (Kyoto).
Oct. 17, 2018
Fangfang Wang :
4 p.m. in 636 SEO
Abstract
In this talk, a new parameter-driven model for multivariate time series of counts is discussed. The time series is not necessarily stationary. The mean process is modelled as the product of modulating factors and unobserved stationary processes. The former characterizes the long-run movement in the data, while the latter is responsible for rapid fluctuations and other unknown or unavailable covariates. The unobserved stationary processes evolve independently of the past observed counts, and might interact with each other. We express the multivariate unobserved stationary processes as a linear combination of possibly low-dimensional factors that govern the contemporaneous and serial correlation within and across the observed counts. Regression coefficients in the modulating factors are estimated via pseudo maximum likelihood estimation, and identification of common factor(s) is carried out through eigenanalysis on a positive definite matrix that pertains to the autocovariance of the observed counts at nonzero lags. Theoretical validity of the two-step estimation procedure is presented. We also provide numerical results that corroborate the theoretical findings. Finally, we illustrate the use of the proposed model through an application to the numbers of National Science Foundation funding awarded to seven research universities from January 2001 to December 2012.
Oct. 24, 2018
Mladen Kolar :
4 p.m. in 636 SEO
Abstract
We present a recent line of work on estimating differential networks and conducting statistical inference about parameters in a high-dimensional setting. First, we consider a Gaussian setting and show how to directly learn the difference between the graph structures. A debiasing procedure will be presented for construction of an asymptotically normal estimator of the difference. Next, building on the first part, we show how to learn the difference between two graphical models with latent variables. Linear convergence rate is established for an alternating gradient descent procedure with correct initialization. Simulation studies illustrate performance of the procedure. We also illustrate the procedure on an application in neuroscience. Finally, we will discuss how to do statistical inference on the differential networks when data are not Gaussian.
Oct. 31, 2018
Yi-Lin Chiu :
4 p.m. in 636 SEO
Abstract
This presentation addresses the basic science for clinical pharmacology: effective and safe drug administration. We use basic mathematics and statistics to introduce the applications to pharmacology, including characterizing the drug concentration, dose selection, and dosing strategy. The utility of clinical pharmacology will be explained: We will show how drugs work, rather than asking the audience to memorize information about individual drugs. Therefore, we can understand why drugs are given, as well as when they should be given, and come up with a better way to improve the effectiveness of drug administrations.
Nov. 7, 2018
Annie Qu :
4 p.m. in 636 SEO
Abstract
Recommender systems have been widely adopted by electronic commerce and entertainment industries for individualized prediction and recommendation, which benefit consumers and improve business intelligence. In this article, we propose an innovative method, namely the recommendation engine of multilayers (REM), for tensor recommender systems. The proposed method utilizes the structure of a tensor response to integrate information from multiple modes, and creates an additional layer of nested latent factors to accommodate between-subjects dependency. One major advantage is that the proposed method is able to address the “cold-start" issue in the absence of information from new customers, new products or new contexts. Specifically, it provides more effective recommendations through sub-group information. To achieve scalable computation, we develop a new algorithm for the proposed method, which incorporates a maximum block improvement strategy into the cyclic block-wise-coordinate-descent algorithm. In theory, we investigate both algorithmic properties for global and local convergence, along with the asymptotic consistency of estimated parameters. Finally, the proposed method is applied in simulations and IRI marketing data with 116 million observations of product sales. Numerical studies demonstrate that the proposed method outperforms existing competitors in the literature. This is joint work with Xuan Bi and Xiaotong Shen.
Nov. 14, 2018
Renming Song :
4 p.m. in 636 SEO
Abstract
In this talk I will discuss heat kernel estimates for critical perturbations
of non-local operators. To be more precise, let $X$ be the reflected
$\alpha$-stable process in the closure of a smooth open set $D$, and
$X^D$ the process killed upon exiting $D$. We consider potentials of the
form $\kappa(x)=C\delta_D(x)^{-\alpha}$ with positive $C$ and the
corresponding Feynman-Kac semigroups. Such potentials do not belong
to the Kato class. We obtain sharp two-sided estimates for the heat
kernel of the perturbed semigroups. The interior estimates of the
heat kernels have the usual $\alpha$-stable form, while the boundary
decay is of the form $\delta_D(x)^p$ with non-negative
$p\in [\alpha-1, \alpha)$ depending on the precise value of the
constant $C$. Our result recovers the heat kernel estimates of both
the censored and the killed stable process in $D$. Analogous
estimates are obtained for the heat kernel of the Feynman-Kac
semigroup of the $\alpha$-stable process in
${\mathbf R}^d\setminus \{0\}$ through the potential $C|x|^{-\alpha}$.
All estimates are derived from a more general result described as follows:
Let $X$ be a Hunt process on a locally compact separable metric space in
a strong duality with $\widehat{X}$. Assume that transition densities of
$X$ and $\widehat{X}$ are comparable to the function $\widetilde{q}(t,x,y)$
defined in terms of the volume of balls and a certain scaling function.
For an open set $D$ consider the killed process $X^D$, and a critical
smooth measure on $D$ with the corresponding positive additive functional
$(A_t)$. We show that the heat kernel of the the Feynman-Kac semigroup
of $X^D$ through the multiplicative functional $\exp(-A_t)$ admits the
factorization of the form
${\mathbf P}_x(\zeta >t)\widehat{\mathbf P}_y(\widehat{\zeta}>t)\widetilde{q}(t,x,y)$.
This is joint work with Soobin Cho, Panki Kim and Zoran Vondracek.
Nov. 28, 2018
Linyi Zhang :
4 p.m. in 636 SEO
Feb. 20, 2019
Subhashis Ghoshal :
4 p.m. in 636 SEO
Abstract
The filament of a smooth function f consists of local maximizers of f when moving in a certain direction. The filament is an important geometrical feature of the surface of the graph of a function. It is also considered as an important lower dimensional summary in analyzing multivariate data. There have been some recent theoretical studies on estimating filaments of a density function using a nonparametric kernel density estimator. In this talk, we consider a Bayesian approach and concentrate on the nonparametric regression problem. We study the posterior contraction rates for filaments using a finite random series of B-splines prior on the regression function. Compared with the kernel method, this has the advantage that the bias can be better controlled when the function is smoother, which allows obtaining better rates. Under an isotropic Holder smoothness condition, we obtain the posterior contraction rate for the filament under two different metrics --- a distance of separation along an integral curve, and the Hausdorff distance between sets. Moreover, we construct credible sets of optimal size for the filament with sufficient frequentist coverage. We study the performance of our proposed method through a simulation study and apply on a dataset on California earthquakes to assess the fault-line of the maximum local earthquake intensity.
Based on joint work with my former graduate student, Dr. Wei Li, Assistant Professor, Syracuse University, New York.
Feb. 27, 2019
William Li :
4 p.m. in 636 SEO
Abstract
While literature on constructing efficient experimental designs has been plentiful, how best to incorporate prior information when assigning factors to the columns has received little attention. This talk summarizes a series of recent studies that focus on information of individual columns. For regular designs, we propose the individual word length pattern (iWLP) that can be used to rank columns. With prior information on how likely a factor is important, iWLP can be used to intelligently assign factors to columns, and select the best designs to accommodate such prior information. This criterion is then extended to study nonregular designs, which we denote as the individual generalized word length pattern (iGWLP). We illustrate how iGWLP helps to identify important differences in the aliasing that is likely otherwise missed. Given the complexity of characterizing partial aliasing, iGWLP will help practitioners make more informed assignment of factors to columns when utilizing nonregular fractions. The theoretical justifications of the proposed iGWLP are provided in terms of statistical model and projection properties. In the third part, we consider clear effects involving an individual column (iCE). Motivated by a real application, we introduce the clear effects pattern, derived from iCE, and propose a class of designs called maximized clear effects pattern (MCEP) designs. We compare MCEP designs with commonly used minimum aberration designs and MaxC2 designs that maximize the number of clear two-factor interaction. We also extend the definition of iCE and MCEP designs by considering blocking schemes.
March 6, 2019
Yue Niu :
4 p.m. in 636 SEO
Abstract
In many applications such as copy number variant (CNV) detection, the goal is to identify short segments on which the observations have different means or medians from the background. Those segments are usually short and hidden in a long sequence, and hence are very challenging to find. We study a super scalable short segment (4S) detection algorithm in this paper. This nonparametric method clusters the locations where the observations exceed a threshold for segment detection. It is computationally efficient and does not rely on Gaussian noise assumption. Moreover, we develop a framework to assign significance levels for detected segments. We demonstrate the advantages of our proposed method by theoretical, simulation, and real data studies.
March 13, 2019
Hani Aldirawi :
4 p.m. in 636 SEO
Abstract
Modeling sparse and discrete data such as microbiome and insurance claim data is challenging due to the exceeded number of zeros. Many probabilistic models have been used for modeling sparse data, including Poisson, negative binomial, zero-inflated Poisson, and zero-inflated negative binomial models. We propose a statistical procedure for identifying the most appropriate discrete probabilistic models for zero-inflated or Hurdle models based on the p-value of the discrete Kolmogorov-Smirnov (KS) test when the population parameters are unknown. We develop a general procedure for estimating the parameters for a large class of zero-inflated models and Hurdle models. We also develop a general likelihood ratio test based on Neyman-Pearson lemma for choosing the best model when appropriate ones are more than one.
March 20, 2019
Mengjia Yu :
4 p.m. in 636 SEO
Abstract
I will discuss two approaches of the cumulative sum (CUSUM) statistics and the U-statistics in change point problems for high-dimensional location-shift. Both works are non-parametric, fully data-dependent and enjoying strong theoretical guarantees under arbitrary dependence structures.
1. Based on the $\ell^{\infty}$-norm of the CUSUM statistics, we study inference and identification for high-dimensional mean vectors. For the problem of testing existence of a change point in an independent sample generated from the mean-shift model, we introduce a Gaussian multiplier bootstrap to calibrate critical values of the CUSUM test statistics. For the problem of estimating the change point location once it is detected, two estimators are proposed by maximizing the $\ell^{\infty}$-norm of the generalized CUSUM statistics at two different weighting scales. In both problems, dimension impacts the rate of convergence only through the logarithm factors, and therefore consistency of the CUSUM location estimators is possible when $p$ is much larger than $n$.
2. In cases where mean does not exist, we consider signal cancellations in the general U-statistics framework with anti-symmetric kernels of order 2, and proposed another test that is more robust to detect location-shift by selecting bounded kernels. The $\ell^{\infty}$-norm of the U-statistic and its Gaussian multiplier bootstrap approximation are investigated, and no tuning parameter is needed in this scheme. Subject to mild conditions kernels, we derive similar rates of uniform convergence to our CUSUM-based test. Connection of two approaches and numeric studies are also provided.
April 3, 2019
Marco Ferreira :
4 p.m. in 636 SEO
Abstract
We discuss classes of dynamic multiscale models for multivariate Gaussian spatiotemporal data. First, we develop multiscale spatial factorizations to decompose the data at each time point into spatiotemporal multiscale coefficients. We then connect these spatiotemporal multiscale coefficients through time with state-space evolutions. Further, we propose simulation-based Bayesian posterior analysis. In particular, we develop filtering equations for updating of information forward in time and smoothing equations for integration of information backward in time, and use these equations to develop forward filter backward samplers for the spatiotemporal multiscale coefficients. Because the multiscale coefficients are conditionally independent a posteriori, our Bayesian posterior analysis is scalable, computationally efficient, and highly parallelizable. Finally, we illustrate the usefulness of our dynamic multiscale spatiotemporal methodology with applications to multivariate spatiotemporal data on temperatures in the upper troposphere and lower stratosphere over North America.
April 10, 2019
Jane Qian :
4 p.m. in 636 SEO
Abstract
ICH is the International Council for Harmonization of Technical Requirements for Pharmaceuticals for Human Use. ICH E9 guideline (“Statistical Principles for Clinical Trials”) was published in 1995 and has since served as a foundation of regulatory guidance on major statistical aspects of confirmatory clinical trials. In October 2014, the Steering Committee of ICH endorsed the formation of an expert working group to develop an addendum to the ICH E9 guideline. In 2017, the addendum ICH E9(R1) (Estimands and sensitivity analyses in clinical trials) was released for public comment. ICH E9(R1) focus on two topics involving randomized confirmatory clinical trials: estimands and sensitivity analyses. Both topics are motivated by the need to improve the precision with which scientific questions of interest are formulated and addressed by clinical researchers and regulators, specifically in the context of post-randomization (intercurrent) events such as use of rescue medication or missing data. In this seminar, an overview of ICH E9(R1) will be given with the focus on why it was necessary to develop this ICH E9 addendum, what it entails and the impact of the addendum on future clinical trial design and analyses.
April 17, 2019
Yehua Li :
4 p.m. in 636 SEO
Abstract
We consider spatially dependent functional data collected under a geostatistics setting, where locations are sampled from a spatial point process and a random function is observed at each location. The functional response is the sum of a spatially dependent functional effect and a spatially independent functional nugget effect. Observations on each function are made on discrete time points and contaminated with measurement errors. Under the assumption of spatial stationarity and isotropy, we propose a tensor product spline estimator for the spatio-temporal covariance function. If a coregionalization covariance structure is further assumed, we propose a new functional principal component analysis method that borrows information from neighboring functions. Under a unified framework for both sparse and dense functional data, where the number of observations per curve is allowed to be of any rate relative to the number of functions, we develop the asymptotic convergence rates for the proposed estimators. Advantages of the proposed approach over existing methods are demonstrated through simulation studies and a real data application to the home price-rent ratio data in the San Francisco Bay Area.
Junhui Wang :
3 p.m. in 636 SEO
Abstract
In recent years, there has been a growing demand to develop efficient recommender systems which track users' preferences and recommend potential
items of interest to users. In this talk, I will present a smooth collaborative
recommender system to utilize dependency information among users and
items which share similar characteristics under the singular value decomposition framework. The proposed method incorporates the neighborhood
structure among user-item pairs by exploiting covariates to improve the prediction performance. One key advantage of the proposed method is that it
leads to more efficient recommendation for "cold-start" users and items,
whose preference information is completely missing from the training set.
As this type of data involves large-scale customer records, efficient scheme
will be proposed to achieve scalable computing. The advantage is confirmed
in a variety of simulated experiments as well as one large-scale real example
on <i>Last.fm</i> music listening counts. If time permits, the asymptotic properties
will also be discussed.
April 24, 2019
Wenhui Sheng :
4 p.m. in 636 SEO
Abstract
We propose a new sufficient dimension folding method using distance covariance for regression in which the predictors are matrix- or array-valued. The method works efficiently without strict assumptions on the predictor. It is modelfree and neither smoothing techniques or selection of tuning parameters is needed. Moreover, it works for both univariate and multivariate response cases. We use two approaches to estimate the structural dimensions: bootstrap method and a new method of local search. Simulations and real data analysis support the efficiency and effectiveness of the method.
May 1, 2019
Solomon Harrar :
4 p.m. in 636 SEO
Abstract
Recent results for high-dimensional inference make assumptions that require weak dependence (pseudo independence) between the variables. These requirements fail to be satisfied, for example, for all elliptically contoured distributions except for normal distribution. In this talk, we present weaker dependence conditions for high-dimensional asymptotic theory. With these conditions the scope of application of many high-dimensional results broadens substantially. For example, mixing-type dependence and general conditions on variance of quadratic forms are covered. The application of the new conditions will be demonstrated with high-dimensional tests for comparing group differences in terms of means and in terms of Mann-Whitney effects. The later is particularly useful for non-metric data such as ordered categorical data, and also for skewed and heavy tailed continuous data. Simulation results show favorable performance of these tests. Data from Electroencephalograph (EEG) experiment is analyzed to illustrate these applications.
The results presented in this talk are joint works with Xiaoli Kong, Department of Mathematics and Statistics, Loyola University-Chicago
May 13, 2019
Abhyuday Mandal :
3 p.m. in 636 SEO
Abstract
Computer experiments with both quantitative and qualitative inputs are commonly used in science and engineering applications. Constructing desirable emulators for such computer experiments remains a challenging problem. Here we propose an easy-to-interpret Gaussian process (EzGP) model for computer experiments to reflect the change of the computer model under different level combinations of qualitative factors. The proposed modeling strategy, based on an additive Gaussian process, is flexible to address the heterogeneity of computer models involving multiple qualitative factors. We also develop two useful variants of the EzGP model to achieve computation efficiency when dealing with high dimensional data and large data size. The merits of these models are illustrated by a real data application and several numerical examples.
William Li :
4 p.m. in 636 SEO
Abstract
Exchange-type of algorithms have been commonly used in design construction problems. In recent years, algorithms based on Particle Swarm Optimization (PSO) techniques have been proposed to construct optimal designs. PSO algorithms have been developed mostly for continuous-type of problems. In this talk we develop a general class of Particle Swarm Exchange algorithms, targeting discrete-type of design problems . We used the proposed algorithm to construct a class of optimal model-discriminating designs. It is shown that the algorithms work both efficiently and effectively. The algorithm compared favorably with the coordinate-exchange algorithm - one of the most commonly used algorithms. And we obtained model-discriminating designs that are comparable and sometime better than existing results.
Aug. 28, 2019
Yichao Wu :
4 p.m. in 636 SEO
Sept. 4, 2019
Zhihua Su :
4 p.m. in 636 SEO
Abstract
Sparse partial least squares (SPLS) is widely used in applied sciences as a method that performs dimension reduction and variable selection simultaneously in linear regression. Several implementations of SPLS have been derived, among which the SPLS proposed in Chun and Keleş (2010) is very popular and highly cited. However, for all of these implementations, the theoretical properties of SPLS are largely unknown. In this paper, we propose a new version of SPLS, called the envelope-based SPLS, using a connection between envelope models and partial least squares (PLS). We establish the consistency, oracle property and asymptotic normality of the envelope-based SPLS estimator. The large-sample scenario and high-dimensional scenario are both considered. We also develop the envelope-based SPLS estimators under the context of generalized linear models, and discuss its theoretical properties including consistency, oracle property and asymptotic distribution. Numerical experiments and examples show that the envelope-based SPLS estimator has better variable selection and prediction performance over the existing SPLS estimators.
Sept. 11, 2019
Stacey Tannenbaum :
4 p.m. in 636 SEO
Abstract
Pharmacometrics is the application of biological and pharmacological science and statistical/ mathematical/ computational methods to optimize pharmaceutical development. The goal of a pharmacometrician is to get the right dose of the right drug to the right patient at the right time! Pharmacometrics includes a wide span of models and applications, but the primary focus of the seminar is on understanding the pharmacokinetics of the drug (how a drug is absorbed, processed, distributed, and eliminated, and how much is in the plasma and site of action at a given time). Every subject will have their own unique concentration-time profile for a given dose and formulation, which is dependent upon intrinsic and extrinsic factors (weight, smoking status, other drugs, health state, etc). One important job of the pharmacometrician is to understand the quantitative impact of these factors on the pharmacokinetics, and to determine which have enough of an impact to make a change to the recommended dose. Once a model is fit to a patient population, simulations can be performed to assess the outcome with different inputs (higher and lower doses, less/more frequent doses, etc). The seminar will also include a discussion of some of the technical details of Pharmacometrics, including some of the mathematical methods, software, data sources, validation techniques, and challenges that pharmacometricians face in their day-to-day work.
Sept. 18, 2019
Jun Li :
4 p.m. in 636 SEO
Abstract
Clustering analysis, in its traditional setting, identifies groupings of samples from a single population/condition. We consider a different setting when the data available are samples from two different conditions, such as cells before and after drug treatment. Cell types in cell populations change as the condition changes: some cell types die out, new cell types may emerge, and surviving cell types evolve to adapt to the new condition. Using single-cell RNA-sequencing data that measure the gene expression of cells before and after the condition change, we propose an algorithm, SparseDC, which identifies cell types, traces their changes across conditions, and identifies genes which are marker genes for these changes. By solving a unified optimization problem, SparseDC completes all three tasks simultaneously. As a general algorithm that detects shared/distinct clusters for two groups of samples, SparseDC can be applied to problems outside the field of biology.
Sept. 25, 2019
Dan Nettleton :
4 p.m. in 636 SEO
Abstract
Breiman's seminal paper on random forests has more than 30,000 citations according to Google Scholar. The impact of Breiman's random forests on machine learning, data analysis, data science, and science in general is difficult to measure but unquestionably substantial. The virtues of random forest methodology include no need to specify functional forms relating predictors to a response variable, capable performance for low-sample-size high-dimensional data, general prediction accuracy, easy parallelization, few tuning parameters, and applicability to a wide range of prediction problems with categorical or continuous responses. Like many algorithmic approaches to prediction, random forests are typically used to produce point predictions that are not accompanied by information about how far those predictions may be from true response values. From the statistical point of view, this is unacceptable; a key characteristic that distinguishes statistically rigorous approaches to prediction from others is the ability to provide quantifiably accurate assessments of prediction error from the same data used to generate point predictions. Thus, we develop a prediction interval -- based on a random forest prediction -- that gives a range of values that will contain an unknown continuous univariate response with any specified level of confidence. We illustrate our proposed approach to interval construction with examples and demonstrate its effectiveness relative to other approaches for interval construction using random forests.
Oct. 2, 2019
Ning Hao :
4 p.m. in 636 SEO
Abstract
The variance of noise plays an important role in many change-point detection tools and their inference. For example, in binary segmentation or other stepwise detection methods, the variance is necessary to decide when to stop the procedure. In practice, people usually use some ad-hoc methods to estimate the noise variance. However, these methods may be problematic when there are many change points. We will introduce an equivariant variance estimator and show its advantages over existing methods. This talk is based on a joint work with Yue S. Niu and Han Xiao.
Oct. 9, 2019
Jianfeng Zhang :
3 p.m. in 636 SEO
Abstract
Motivated by option pricing in a financial market with rough volatility, we study backward SDEs in a framework where the (forward) state process satisfies a Volterra type SDE, with fractional Brownian motion as a typical example. Such processes are neither Markov processes nor semimartingales, and most notably, they feature a certain time inconsistency which makes any direct application of Markovian ideas impossible without passing to a path-dependent framework. Our main result is a functional Ito formula, extending the seminal work of Dupire to our more general framework. In particular, unlike in Dupire's setting where one needs only to consider the stopped paths, here we need to concatenate the observed path up to the current time with a certain smooth observable curve derived from the distribution of the future paths. This new feature is due to the time inconsistency involved in this paper. We then derive the path dependent PDEs for the backward problems. The talk is based on a joint work with Frederi Viens.
Hira Koul :
4 p.m. in 636 SEO
Abstract
We develop analogs of the two classes of weighted empirical min- imum distance estimators of the underlying parameters in linear and nonlinear regression models when covariates are observed with Berk- son measurement error. One class is based on the integral of the square of symmetrized weighted empirical of residuals while the other is based on a similar integral involving a weighted empirical of residual ranks. The former class requires the regression and measurement errors to be symmetric around zero while the latter class does not need any such assumption. The first class of estimators includes the analogs of the least absolute deviation and Hodges-Lehmann estimators while the second class includes an estimator that is asymptotically more effi- cient than these two estimators at some error distributions when there is no measurement error. In the case of linear model, no knowledge of the measurement error distribution is needed. Such information is typically needed for non-linear models. We first develop these esti- mators for nonlinear models when the measurement error distribution is known and then their analogs, when this distribution is not known but validation data is available.
Oct. 16, 2019
Yixin Fang :
4 p.m. in 636 SEO
Abstract
Abstract: Randomized controlled clinical trials (RCTs) are the gold standard for evaluating the safety and efficacy of pharmaceutical drugs, but in many cases their costs, duration, limited generalizability, and ethical or technical feasibility have caused some to look for real-world studies as alternatives. On the other hand, real-world data may be much less convincing due to the lack of randomization and the presence of confounding bias. In this article, we propose a statistical roadmap to translate real-world data (RWD) to robust real-world evidence (RWE). The Food and Drug Administration (FDA) is working on guidelines, with a target to release a draft by 2021, to harmonize RWD applications and monitor the safety and effectiveness of pharmaceutical drugs using RWE. The proposed roadmap aligns with the newly released framework for FDA's RWE Program in December 2018 and we hope this statistical roadmap is useful for statisticians who are eager to embark on their journeys in the real-world research.
Oct. 23, 2019
Xianyang Zhang :
4 p.m. in 636 SEO
Abstract
We present new metrics to quantify and test for (i) the equality of distributions and (ii) the independence between two high-dimensional random vectors. We show that the energy distance based on the usual Euclidean distance cannot completely characterize the homogeneity of two high-dimensional distributions in the sense that it only detects the equality of means and the traces of covariance matrices in the high-dimensional setup. We propose a new class of metrics which inherit the desirable properties of the energy distance/distance covariance in the low-dimensional setting and is capable of detecting the homogeneity of/ completely characterizing independence between the low-dimensional marginal distributions in the high dimensional setup. We further propose t-tests based on the new metrics to perform high-dimensional two-sample testing/ independence testing and study its asymptotic behavior under both high dimension low sample size (HDLSS) and high dimension medium sample size (HDMSS) setups. The computational complexity of the t-tests only grows linearly with the dimension and thus is scalable to very high dimensional data. We demonstrate the superior power behavior of the proposed tests for homogeneity of distributions and independence via both simulated and real datasets.
Oct. 30, 2019
Minge Xie :
4 p.m. in 636 SEO
Abstract
This paper proposes a new and effective simulation-based approach, called Repro Sampling method, to conduct statistical inference in high dimensional linear models. The Repro method creates and studies the performance of artificial samples (referred to as Repro samples) that are generated by mimicking the sampling mechanism that generated the true observed sample. By doing so, this method provides a new way to quantify model and parameter uncertainty and provide confidence sets with guaranteed coverage rates on a wide range of problems. A general theoretical framework and an effective Monte-Carlo algorithm, with supporting theories, are developed for high dimensional linear models. This method is used to jointly create confidence sets of selected models and model coefficients, with both exact and asymptotic inferences theories provided. It also provides a theoretical development to support the computational efficiency. Furthermore, this development allows us to handle inference problems involving covariates that are perfectly correlated. A new and intuitive graphical tool to present uncertainties in model selection and regression parameter estimation is also developed. We provide numerical studies to demonstrate the utility of the proposed method in a range of problems. Numerical comparisons suggest that the method is far better (in terms of improved coverage rates and significantly reduced sizes of confidence sets) than the approaches that are currently used in the literature. The development provides a simple and effective solution for the difficult post-selection inference problems.
Nov. 6, 2019
Dr. Ching Jin :
4 p.m. in 636 SEO
Abstract
Diffusion processes are central to human interactions. One common prediction of the current modeling frameworks is that
initial spreading dynamics follow exponential growth. Here we find that, for subjects ranging from mobile handsets to automobiles
and from smartphone apps to scientific fields, early growth patterns follow a power law with non-integer exponents.
We test the hypothesis that mechanisms specific to substitution dynamics may play a role, by analyzing unique data tracing
3.6 million individuals substituting different mobile handsets. We uncover three generic ingredients governing substitutions,
allowing us to develop a minimal substitution model, which not only explains the power-law growth, but also collapses diverse
growth trajectories of individual constituents into a single curve. These results offer a mechanistic understanding of power-law
early growth patterns emerging from various domains and demonstrate that substitution dynamics are governed by robust
self-organizing principles that go beyond the particulars of individual systems.
This talk is based on my recent Nature Human Behaviour paper (attached with the email). If we have enough time, I would also like to share a couple of follow-ups of the paper or a couple of related projects we are working on recently.
Nov. 13, 2019
Jennifer Pajda-Delao :
4 p.m. in 636 SEO
Abstract
This talk will introduce survey sampling along with some sampling designs. Then we discuss the minimum, maximum, and median as important parameters in finite population sampling. We can prove that there are no unbiased estimators of the minimum, maximum, or median for finite population sampling under any sampling design except census. We then identify and characterize a family of sampling designs such that, under these designs, the sample median is a median-unbiased estimator of the population median. In particular, we consider the simple random sampling case.
Nov. 20, 2019
Xiangrong Yin :
4 p.m. in 636 SEO
Abstract
The T-central subspace, introduced by Luo, Li and Yin (2014), allows one to perform sufficient dimension reduction for
any statistical functional of interest. We propose a general estimator using (third) moment kernel to estimate the T-central subspace.
In this talk, we particularly focus on central mean subspace via the regression mean function, and central subspace via Fourier
transform or slicing. Theoretical results are established and simulation studies show the advantages of our proposed methods.
Dec. 4, 2019
Seonghyun Jeong :
4:15 p.m. in 636 SEO
Abstract
This study investigates frequentist properties of Bayesian high-dimensional logit models for categorical response variables. For high-dimensional regression coefficients, group sparse modeling is adopted to handle model selection with categorical responses. A product of a point mass and a Laplace-type distribution is used for the prior distribution on sparse regression coefficients. The procedure exhibits nearly optimal posterior contraction. A shape approximation to the posterior distribution is characterized to show model selection consistency. The distributional approximation also leads to a Bernstein-von Mises theorem for uncertainty quantification through credible sets with guaranteed frequentist coverage.
Feb. 19, 2020
Trambak Banerjee :
4:15 p.m. in 636 SEO
Abstract
We develop a novel shrinkage rule for prediction in a high-dimensional non-exchangeable hierarchical Gaussian model with an unknown spiked covariance structure. We propose a family of commutative priors for the mean parameter, governed by a power hyper-parameter, which encompasses from perfect independence to highly dependent scenarios. Corresponding to popular loss functions such as quadratic, generalized absolute, and linex losses, these prior models induce a wide class of shrinkage predictors that involve quadratic forms of smooth functions of the unknown covariance. By using uniformly consistent estimators of these quadratic forms, we propose an efficient procedure for evaluating these predictors which outperforms factor model based direct plug-in approaches. We further improve our predictors by introspecting possible reduction in their variability through a novel coordinate-wise shrinkage policy that only uses covariance level information and can be adaptively tuned using the sample eigen structure. We extend our methodology to aggregation based prescriptive analysis of generic multidimensional linear functionals of the predictors that arise in many contemporary applications involving forecasting decisions on portfolios or combined predictions from dis-aggregative level data. We propose an easy-to-implement functional substitution method for predicting linearly aggregative targets and establish asymptotic optimality of our proposed procedure. We present simulation experiments as well as real data examples illustrating the efficacy of the proposed method.
Feb. 26, 2020
Hyun-Jung Kim :
4 p.m. in 636 SEO
Abstract
In this talk, we discuss recent discoveries in statistical inference for stochastic partial differential equations (SPDEs). We mainly focus on parameter estimation problems in stochastic evolution equations driven by additive noise: 1. space-time and 2. space-only colored (or white) noise. The goal of this talk is to derive "good" estimators in the sense that they are consistent and asymptotically normal to a true parameter in a specific asymptotic regime when continuous or discrete sampling of the solution process is available.
March 4, 2020
Xiaotong Shen :
4 p.m. in 636 SEO
Qi Feng :
3 p.m. in 636 SEO
Abstract
The classical models for asset processes in math finance are SDEs driven by Brownian motion of the following type
$X_t=x+\int_0^tb(s,X_s)ds+\int_0^t\sigma(s,X_s)\circ dB_s$.
Then $u(t,X_t)=\mathbb E[{g(X_T)}|\mathcal F_{t}^X]$ is a deterministic function of $X_t$ and $u(t,x)$ solves a parabolic PDE. The cubature formula is first constructed to numerically compute functionals like $\mathbb E^{\mathbb P}[g(X_T)]$, which can be seen as a discrete approximation of the infinite dimensional Wiener measure (denoted as $\mathbb P$). In this talk, we will consider that the asset process follows a rough volatility model. For example, in the rough Heston model, the process $X_t$ is the solution of Volterra type SDEs. In this case, $X$ itself is non-Markovian, then $u(t,X_t)$ will depend on the whole path of $(X_s)_{0\le s\le t}$ and $u(t,X_{[0,t]})$ solves the so-called Path Dependent PDE (PPDE). We propose a new algorithm to numerically solve PPDE by using cubature type formulas for Volterra SDEs. The cubature formula for Volterra SDEs is solved by using machine learning method. In the end, I will show some numerical examples. The talk is based on a joint work with Jianfeng Zhang.
March 18, 2020
Willaim Li :
4 p.m. in 636 SEO
April 1, 2020
Ivan Nourdin :
4 p.m. in 636 SEO
Abstract
TBA
April 8, 2020
Lanju Zhang :
4 p.m. in 636 SEO
April 15, 2020
Rina Foygel Barber :
3 p.m. in 636 SEO
April 22, 2020
Yongzhao Shao :
4 p.m. in 636 SEO
April 29, 2020
Peng Zeng :
4 p.m. in 636 SEO
Aug. 26, 2020
No speaker :
4 p.m. in Zoom
Sept. 2, 2020
Yuexiao Dong :
4 p.m. in Zoom
Abstract
We introduce a novel framework for model-free variable selection with matrix-valued predictors. To test the importance of rows, columns, and submatrices of the predictor matrix in terms of predicting the response, three types of hypotheses are formulated under a unified framework. The asymptotic properties of the test statistics under the null hypothesis are established and a permutation testing algorithm is also introduced to approximate the distribution of the test statistics. A maximum ratio criterion (MRC) is proposed to facilitate the model-free variable selection. Unlike the traditional stepwise regression procedures that require calculating p-values at each step, the MRC is a non-iterative procedure that does not require p-value calculation and is guaranteed to achieve variable selection consistency under mild conditions. Performance of the proposed method is evaluated in extensive simulations and demonstrated through the analysis of an electroencephalography data.
Sept. 16, 2020
Xinran Li :
4 p.m. in 636 SEO
Abstract
Randomization (a.k.a. permutation) inference is typically interpreted as testing Fisher's ``sharp'' null hypothesis that all effects are exactly zero. This hypothesis is often criticized as uninteresting and implausible. We show, however, that many randomization tests are also valid for a ``bounded'' null hypothesis under which effects are all negative (or positive) for all units but otherwise heterogeneous. The bounded null is closely related to important concepts such as monotonicity and Pareto efficiency. Inverting tests of this hypothesis yields confidence intervals for the maximum (or minimum) individual treatment effect. We then extend randomization tests to infer other quantiles of individual effects, which equivalently infers proportions of units with effects larger (or smaller) than any thresholds. The proposed confidence intervals for all quantiles of individual effects are simultaneously valid, in the sense that no correction due to multiple analyses is needed. In sum, we provide a broader justification for Fisher randomization tests, and develop exact nonparametric inference for quantiles of heterogeneous individual effects. The proposed methods move beyond usual constant effects under Fisher randomization tests and average effect in Neyman's repeated sampling inference. We illustrate our methods with simulations and applications, where we find that Stephenson rank statistics often provide the most informative results.
Sept. 23, 2020
Rina Foygel Barber :
4 p.m. in Zoom
Abstract
Goodness-of-fit (GoF) testing is ubiquitous in statistics, with direct ties to model selection, confidence interval construction, conditional independence testing, and multiple testing, just to name a few applications. While testing the GoF of a simple (point) null hypothesis provides an analyst great flexibility in the choice of test statistic while still ensuring validity, most GoF tests for composite null hypotheses are far more constrained, as the test statistic must have a tractable distribution over the entire null model space. A notable exception is co-sufficient sampling (CSS): resampling the data conditional on a sufficient statistic for the null model guarantees valid GoF testing using any test statistic the analyst chooses. But CSS testing requires the null model to have a compact (in an information-theoretic sense) sufficient statistic, which only holds for a very limited class of models; even for a null model as simple as logistic regression, CSS testing is powerless. In this paper, we leverage the concept of approximate sufficiency to generalize CSS testing to essentially any parametric model with an asymptotically-efficient estimator; we call our extension “approximate CSS” (aCSS) testing. We quantify the finite-sample Type I error inflation of aCSS testing and show that it is vanishing under standard maximum likelihood asymptotics, for any choice of test statistic. We apply our proposed procedure both theoretically and in simulation to a number of models of interest to demonstrate its finite-sample Type I error and power.
This work is joint with Lucas Janson.
Sept. 30, 2020
Jisu Kim :
4 p.m. in Zoom
Abstract
Geometric and topological structures can aid statistics in several ways. In high dimensional statistics, geometric structures can be used to reduce dimensionality. High dimensional data entails the curse of dimensionality, which can be avoided if there are low dimensional geometric structures. On the other hand, geometric and topological structures also provide useful information. Structures may carry scientific meaning about the data and can be used as features to enhance supervised or unsupervised learning. In this talk, I will explore how statistical inference can be done on geometric and topological structures. First, given a manifold assumption, I will explore the minimax rates of dimension estimator and reach estimator. First, given a manifold assumption, I will explore the minimax rate for estimating the dimension of the manifold. Second, also under the manifold assumption, I will explore the minimax rate for estimating the reach, which is a regularity quantity depicting how a manifold is smooth and far from self-intersecting. Third, I will investigate inference on cluster trees, which is a hierarchy tree of high-density clusters of a density function. Fourth, I will investigate inference on persistent homology of a density function, which is a representation of topological features of the density function at different levels. Third, I will present R package TDA for computing topological data analysis, which is a set of data analysis tools utilizing topology and includes persistent homology.
Oct. 7, 2020
Qingshuo Song :
4 p.m. in Zoom
Abstract
The characterization of the efficient frontier in Markowitz portfolio optimization is to minimize a linear combination of mean and variance of the terminal stock price. Such a problem is known as the time-inconsistent optimization and the main difficulty is due to the failure of the dynamic programming principle. The existing approaches are game-theoretic framework and decoupling techniques on its FBSDE formulation. In this talk, we will discuss an alternative approach. The key observation is to identify the linear-quadratic structure of the underlying optimization as a function of probability distribution. This leads to explicit solutions of a class of master equations, which provides the optimal strategy to a class of time-inconsistent optimizations. Some extensions to partially observed systems will be considered briefly if time is permitted. The discussion is based on a manuscript available at https://arxiv.org/pdf/1910.05236.pdf.
Oct. 14, 2020
Lynna Chu :
4 p.m. in Zoom
Abstract
We present a new framework for the testing and estimation of change-points, locations where the distribution abruptly changes. While the change-point problem has been extensively studied for low-dimensional data, advances in data collection technology have produced data sequences of increasing volume and complexity. Motivated by the challenges of modern data, we study a non-parametric framework that utilizes similarity information among observations and can be applied to various data types as long as an informative similarity measure on the sample space can be defined. Analytical p-value approximations are also provided, making the methods easy-off-the-shelf tools for real applications.
Oct. 21, 2020
Yuan Ke :
4 p.m. in Zoom
Abstract
We proposes a model-free and data-adaptive feature screening method for ultra-high dimensional data. The proposed method is based on the projection correlation which measures the dependence between two random vectors. This projection correlation based method does not require specifying a regression model, and applies to data in the presence of heavy tails and multivariate responses. It enjoys both sure screening and rank consistency properties under weak assumptions. A two-step approach, with the help of knockoff features, is advocated to specify the threshold for feature screening such that the false discovery rate (FDR) is controlled under a pre-specified level. The proposed two-step approach enjoys both sure screening and FDR control simultaneously if the pre-specified FDR level is greater or equal to 1/s, where s is the number of active features. The superior empirical performance of the proposed method is illustrated by simulation examples and real data applications.
Oct. 28, 2020
Irina Gaynanova :
4 p.m. in Zoom
Abstract
A great number of multivariate statistical methods, such as principal component analysis, discriminant analysis, canonical correlation analysis and graphical lasso to name a few, require the estimate of covariance or correlation matrix of variables as one of the inputs. It is typical to use Pearson sample correlation matrix, which works well at capturing dependencies between normally distributed variables. In this work we consider the problem of estimating dependencies between zero-inflated measurements, which arise in miRNA data, microbiome data, physical activity data, etc. We propose truncated latent Gaussian copula to model the data with excess zeroes, which allows us to derive a rank-based estimator of latent correlation matrix without the estimation of marginal transformation functions. The new methodology is applied for the analysis of associations between gene expression and microRNA data of breast cancer patients, and for inferring the conditional independence graph in quantitate gut microbiome data.
Nov. 4, 2020
Jun Song :
4 p.m. in Zoom
Abstract
In this talk, a general theory and estimation methods for functional linear sufficient dimension reduction will be presented, where both the predictor and the response can be random functions or even vectors of functions. Unlike the existing dimension reduction methods, our approach does not rely on the estimation of conditional mean and conditional variance. Instead, it is based on a new statistical construction --the weak conditional expectation, which is based on Carleman operators and their inducing functions. Weak conditional expectation is a generalization of conditional expectation. Its key advantage is to replace the projection on to an L2-space -- which defines conditional expectation -- by projection on to an arbitrary Hilbert space, while still maintaining the unbiasedness of the related dimension reduction methods. This flexibility is particularly important for functional data, because attempting to estimate a full-fledged conditional mean or conditional variance by slicing or smoothing over the space of vector-valued functions may be inefficient due to the curse of dimensionality. We evaluated the performances of our new methods by simulation and in several applied settings.
Nov. 11, 2020
Yang Feng :
4 p.m. in Zoom
Abstract
We propose a new model-free ensemble classification
framework, Random Subspace Ensemble (RaSE), for sparse classification.
In the RaSE algorithm, we aggregate many weak learners, where each
weak learner is a base classifier trained in a subspace optimally
selected from a collection of random subspaces. To conduct subspace
selection, we propose a new criterion, ratio information criterion
(RIC), based on weighted Kullback-Leibler divergences. The theoretical
analysis includes the risk and Monte-Carlo variance of RaSE
classifier, establishing the weak consistency of RIC, and providing an
upper bound for the misclassification rate of RaSE classifier. An
array of simulations under various models and real-data applications
demonstrate the effectiveness of the RaSE classifier in terms of low
misclassification rate and accurate feature ranking. The RaSE
algorithm is implemented in the R package RaSEn on CRAN. This is joint
work with Ye Tian.
Nov. 18, 2020
Peng Zeng :
4 p.m. in Zoom
Abstract
Huber regression utilizes the Huber loss instead of the common squared loss to achieve the robustness against outliers. It can be regarded as somewhere in the middle of least squares estimate and least absolute deviation. In this talk, we discuss a family of regularized Huber regression models for simultaneous model fitting and variable selection. The prior domain knowledge can be incorporated as linear constraints on parameters. The number of degrees of freedom is a measure of the effective number of parameters used to fit a regression model. It has been used in information criteria for model selection. We derive a formula for the number of degrees of freedom for regularized Huber regression with linear constraints. Simulation studies and real examples are used to demonstrate the application and performance of the proposed methods.
Nov. 25, 2020
Jeong Min Jeon :
4 p.m. in Zoom
Abstract
Analyzing non-Euclidean data is becoming an important topic in modern statistics, as various non-Euclidean data are emerging. However, it is not transparent how one can analyze such non-Euclidean data in many subject areas. In this talk, we introduce a general regression method for analyzing many types of non-Euclidean data. In particular, we consider additive models with some metric-space-valued predictors and Hilbertian responses. The predictors in our setting cover any finite-dimensional-Hilbert-space-valued predictors and Riemannian-manifold-valued predictors. Hence, they allow for Euclidean, compositional, circular, spherical and shape-valued predictors. The response setting is broad as well covering Euclidean, compositional, functional and density-valued responses. We present several real data analysis which show the wide applications of our method. We also present its asymptotic theory.
Dec. 2, 2020
Daren Wang :
4 p.m. in Zoom
Abstract
We consider a general functional regression model, allowing for both functional and high-dimensional vector predictors. Based on this general setting, we propose a penalized least squares estimator in reproducing kernel Hilbert spaces (RKHS), where the penalties enforce both smoothness and sparsity on the functional estimator. We also show that the excess prediction risk of our estimator is minimax optimal under this general model setting. Our analysis reveals an interesting phase transition phenomenon and the optimal excess risk is determined jointly by the sparsity and the smoothness of the functional regression coefficients.
Jan. 13, 2021
Anru Zhang :
4 p.m. in Zoom
Abstract
The analysis of tensor data has become an active research topic in this area of big data. Datasets in the form of tensors, or high-order matrices, arise from a wide range of applications, such as financial econometrics, genomics, and material science. In addition, tensor methods provide unique perspectives and solutions to many high-dimensional problems, such as topic modeling and high-order interaction pursuit, where the observations are not necessarily tensors. High-dimensional tensor problems generally possess distinct characteristics that pose unprecedented challenges to the data science community. There is a clear need to develop new methods, efficient algorithms, and fundamental theory to analyze the high-dimensional tensor data.
In this talk, we discuss some recent advances in high-dimensional tensor data analysis through the consideration of several fundamental and interrelated problems, including tensor SVD and tensor regression. We illustrate how we develop new statistically optimal methods and computationally efficient algorithms that exploit useful information from high-dimensional tensor data based on the modern theories of computation, high-dimensional statistics, and non-convex optimization. Through tensor SVD, we are able to achieve good performance in the denoising of 4D scanning transmission electron microscopy images. Using tensor regression, we are able to use MRI images for the prediction of attention-deficit/hyperactivity disorder.
Jan. 27, 2021
Dr. Tao Liu :
4 p.m. in Zoom
Abstract
In this seminar, the speaker will present the current practices and trends in data analysis in psychology studies. Research on psychology research methods have found that recent empirical studies published in psychology journals are employing more varied and advanced statistical techniques than were employed previously. The most prevalent statistical analysis methods will be presented, along with the trend of data analysis methods in the past few decades. The presenter will use clinical mental health studies to illustrate such changes, including recent studies during COVID-19 pandemic that investigated the impacts of the public health crisis on psychological wellbeing and social attitudes. Presenter will also provide information of public resources for mental health services and self-care.
Feb. 3, 2021
Zhengling Qi :
4 p.m. in Zoom
Abstract
In this talk, I will discuss the batch (off-line) reinforcement learning problem in infinite horizon Markov Decision Processes. Motivated by mobile health applications, we focus on learning a policy that maximizes the long-term average reward. Given limited pre-collected data, we propose a doubly robust estimator for the average reward and show that it achieves statistical efficiency bound. We then develop an optimization algorithm to compute the optimal policy in a parametrized stochastic policy class. The performance of the estimated policy is measured by the difference between the optimal average reward in the policy class and the average reward of the estimated policy. Under some technical conditions, we establish a strong finite-sample regret guarantee in terms of total decision points, demonstrating that our proposed method can efficiently break the curse of horizon. Finally, the performance of the proposed method is illustrated by simulation studies.
Feb. 10, 2021
Jian Zou :
4 p.m. in Zoom
Abstract
Exploring high frequency transaction level financial data is of considerable interest to researchers
and investors. The extra amount of information contained in high-frequency data and keen
interests in high-frequency finance motivate researchers to study dynamic patterns of comovement over multiple trading days. In this paper, we have developed a series of clustering and
biclustering algorithms based on mutual information for high frequency financial time series. We
examine the co-movement probabilities of selected m-tuples of stocks over multiple trading days
under different metrics. Additionally, we propose a unified framework to describe patterns and
monitor the structure of high-dimensional daily or weekly time series that track linkages between
any given m-tuple of stocks over a long time period.
Feb. 17, 2021
Paromita Dubey :
4 p.m. in Zoom
Abstract
In recent years, samples of time-varying object data such as time-varying networks that are not in a vector space have been increasingly collected. These data can be viewed as elements of a general metric space that lacks local or global linear structure and therefore common approaches that have been used with great success for the analysis of functional data, such as functional principal component analysis, cannot be applied directly.
In this talk, I will propose some recent advances along this direction. First, I will discuss ways to obtain dominant modes of variations in time varying object data. I will describe metric covariance, a novel association measure for paired object data lying in a metric space (\Omega d) that we use to define a metric auto-covariance function for a sample of random \Omega -valued curves, where \Omega generally will not have a vector space or manifold structure. The proposed metric auto-covariance function is non-negative definite when the squared metric d^2 is of negative type. Then the eigenfunctions of the linear operator with the auto-covariance function as kernel can be used as building blocks for an object functional principal component analysis for \Omega-valued functional data, including time-varying probability distributions, covariance matrices and time-dynamic networks. Then I will describe how to obtain analogues of functional principal components for time-varying objects by applying Fréchet means and projections of distance functions of the random object trajectories in the directions of the eigenfunctions, leading to real-valued Fréchet scores and object valued Fréchet integrals. This talk is based on joint work with Hans-Georg Müller.
Feb. 24, 2021
Guannan Wang :
4 p.m. in Zoom
Abstract
With the rapid growth of modern technology, many large-scale imaging studies have been or are being conducted to collect massive datasets with large volumes of imaging data, thus boosting the investigation of "next-generation functional data." These enormous collections of imaging data contain interesting information and valuable knowledge, whichhas raised the demand for further advancement in functional data analysis. In this talk, we mainly focus on modeling and inference of the next-generation functional data. We propose using flexible multivariate splines over triangulation or tetrahedral partitions to handle irregular domain of the images that are common in brain imaging studies and in other biomedical imaging applications. The proposed spline estimators are shown to be consistent and asymptotically normal under some regularity conditions. We also provide a computationally efficient estimator of the covariance function and derive its uniform consistency. Finally, we discuss the inferential capabilities of the proposed method. To be more specific, we develop simultaneous confidence corridors for the mean of the next-generation functional data. The procedure is also extended to the two-sample case in which we focus on comparing the mean functions of random samples drawn from two populations. The proposed method is applied to analyze brain Positron Emission Tomography (PET) data of Alzheimer's Disease.
March 3, 2021
Yeonjoo Park :
4 p.m. in Zoom
Abstract
We present a novel spatial model that predicts scalar responses based on functional predictors observed at spatial locations. We incorporate two spatial components in the modeling, (i) spatial correlation between infinite-dimensional functional predictors and (ii) spatially heterogeneous associations between responses and functional covariates at different locations, by introducing a spatially varying functional coefficient model. It allows the functional coefficients to vary with location. To preserve spatial continuity on the low dimensional representation of functional predictors, we employ nonparametric data-adaptive functions for basis expansion under a Bayesian framework and place spatial priors on projection coefficients. We further propose the spatial variable selection, which allows spatially heterogeneous sets of non-null coefficients over locations by borrowing information across neighbors. The basis function estimation, model parameter estimation, and model selection can be jointly performed through Bayesian hierarchical modeling. For the prediction on new observations, we propose the unified approach which enables the estimation of nonparametric basis functions adaptive to new functional predictors and simultaneously draws predictive values from posterior prediction distribution in MCMC implementation. The model performance is demonstrated in simulation studies and an application to a crop yield prediction.
March 10, 2021
Bing Li :
4 p.m. in Zoom
Abstract
We introduce a Sufficient Graphical Model by applying the recently developed nonlinear sufficient dimension reduction techniques to the evaluation of conditional independence. The graphical model is nonparametric in nature, as it does not make distributional assumptions such as the Gaussian or copula Gaussian assumptions. However, unlike a fully nonparametric graphical model, which relies on the high-dimensional kernel to characterize conditional independence, our graphical model is based on conditional independence given a set of sufficient predictors with a substantially reduced dimension. In this way we avoid the curse of dimensionality that comes with a high-dimensional kernel. We develop the population-level properties, convergence rate, and variable selection consistency of our estimate.
By simulation comparisons and an analysis of the DREAM 4 Challenge data set, we demonstrate that our method outperforms the existing methods when the Gaussian or copula Gaussian assumptions are violated, and its performance remains excellent in the high-dimensional setting.
March 17, 2021
Pradeep Singh :
4 p.m. in Zoom
Abstract
The National Cancer Institute and most states keep a cancer data registry so that it can be used by researchers and policy makers to make better healthcare decisions. This data can have missing observations for one or more variables. In particular, the correct stage at diagnosis is sometimes missing from the data due to various reasons. To use the data, different strategies have been used. Researchers often delete individuals from the study who had missing values from even one variable. Another method is to impute the missing values. There are several methods proposed to impute missing values of quantitative variables. But for categorical variables, there have been few methods proposed. Van der Palm, et al. [2016], compared four imputation methods for categorical data. Zhou et al. [2017] has proposed a nonparametric multiple imputation method using the nearest-neighbor approach. This study applied the nonparametric multiple imputation method proposed by Zhou et al. [2017] and a parametric multiple imputation method to lung adenocarcinoma data from the National Cancer Institute. Lung adenocarcinoma is a type of non-small cell lung cancer that typically forms on the outside of the lungs. A Monte Carlo study was done to compare these methods with respect to imputation bias. The study also compared the effect of different levels (10%, 20%, 40%) of missingness, different sizes of the sample, and different fits of the model on these multiple imputation methods.
March 31, 2021
Shan Yu :
4 p.m. in Zoom
Abstract
The estimator of coefficient functions in an functional linear model (FLM) based on a small number of subjects is often inefficient. To address this challenge, we propose an FLM based on fused learning. This talk will describe a sparse multi-group FLM to simultaneously estimate multiple coefficient functions and identify groups, such that coefficient functions are identical within groups and distinct across groups. By borrowing information from relevant subgroups of subjects, our method enhances estimation efficiency while preserving heterogeneity in model parameters and coefficient functions. We use an adaptive fused lasso penalty to shrink coefficient estimates to a common value within each group. To enhance computation efficiency and incorporate neighborhood information, we propose to use graph-constrained adaptive lasso with a highly efficient algorithm. This talk will use two real data examples to illustrate the applications of the proposed method on genotype-by-environment interaction studies.
This talk features joint work with Aaron Kusmec, Lily Wang, and Dan Nettleton.
April 7, 2021
Luo Xiao :
4 p.m. in Zoom
Abstract
Single index models extend standard linear models to account for non-linearity between multivariate predictors and responses.
We study single index models where the unknown coefficients can be formulated as a matrix and enforce
regularization term(s) on the coefficient matrix to induce meaningful structure, e.g., sparsity and low-rank.
We propose an iterative estimation procedure in which an alternating direction method of multipliers (ADMM) algorithm is
employed to accommodate multiple regularization terms.
We focus on two particular models: scalar response on matrix predictor model and multivariate response on multivariate predictor model.
We apply the former model to study nonlinear association between functional connectivity networks and fluid intelligence, and the latter model
to a genetic association study. The work is based on two papers, "Sparse single index models for multivariate responses” which is to appear in
Journal of Computational and Graphical Statistics and “Single index models with functional connectivity network predictors”, which has been tentatively accepted
by Biostatistics.
April 14, 2021
Guan Yu :
4 p.m. in Zoom
Abstract
Weighted nearest neighbor (WNN) classifiers are fundamental non-parametric classifiers for
classification. They have become the methods of choice in many applications where limited
knowledge of the data generation process is available a priori. There exists a vast room of
flexibility in the choice of weights for the neighbors in a WNN classifier. In this talk, I will introduce
a new locally weighted nearest neighbor (LWNN) classifier, which adaptively assigns weights for
different test data points. Given a training data set and a test data point x0, the weights for
classifying x0 in LWNN is obtained by minimizing an upper bound of the conditional expected
estimation error of the regression function at x0. The resultant weights have a neat closed-form
expression, and therefore the computation of LWNN is more efficient than some existing
adaptive WNN classifiers that require estimating the marginal feature density. Like most other
WNN classifiers, LWNN assigns larger weights for closer neighbors. However, in addition to the
ranks of neighbors' distances, the weights in LWNN also depend on the raw values of the
distances. Our theoretical study shows that LWNN achieves the minimax rate of convergence of
the excess risk, when the marginal feature density is bounded away from zero. In the general
case with an additional tail assumption on the marginal feature density, the upper bound of the
excess risk of LWNN matches the minimax lower bound up to a logarithmic term.
April 21, 2021
Xuan Bi :
4 p.m. in Zoom
Abstract
Brain-imaging data have been increasingly used to understand intellectual disabilities. Despite significant progress in biomedical research, the mechanisms for most of the intellectual disabilities remain unknown. Finding the underlying neurological mechanisms has proved difficult, especially in children due to the rapid development of their brains. We investigate verbal reasoning, which is a reliable measure of an individual’s general intellectual abilities, and develop a class of high-order imaging regression models to identify brain subregions which might be associated with this specific intellectual ability. A key novelty of our method is to take advantage of spatial brain structures, and specifically the piecewise smooth nature of most imaging coefficients in the form of high-order tensors. Our approach provides an effective and urgently needed method for identifying brain subregions potentially underlying certain intellectual disabilities. The idea behind our approach is a carefully constructed concept called internal variation (IV). The IV employs tensor decomposition and provides a computationally feasible substitution for total variation, which has been considered suitable to deal with similar problems but may not be scalable to high-order tensor regression. Before applying our method to analyze the real data, we conduct comprehensive simulation studies to demonstrate the validity of our method in imaging signal identification. Next, we present our results from the analysis of a dataset based on the Philadelphia Neurodevelopmental Cohort for which we preprocessed the data including reorienting, bias-field correcting, extracting, normalizing, and registering the magnetic resonance images from 978 individuals. Our analysis identified a subregion across the cingulate cortex and the corpus callosum as being associated with individuals’ verbal reasoning ability, which, to the best of our knowledge, is a novel region that has not been reported in the literature. This finding is useful in further investigation of functional mechanisms for verbal reasoning.
April 28, 2021
Xiaofeng Shao :
4 p.m. in Zoom
Abstract
This talk consists of two parts. In the first part, I will review some
basic idea of self-normalization (SN) for inference of time series in
the context of confidence interval construction and change-point testing in mean.
In the second part, I will present a piecewise linear quantile trend
model to model infection trajectories of COVID-19 daily new cases. To
estimate the change-points in the linear trend, we develop a new segmentation algorithm based on SN test statistics and local scanning.
Data analysis for COVID-19 infection trends in many countries demonstrates the usefulness of our new
model and segmentation method.
Xiaofeng Shao :
4 p.m. in Zoom
Abstract
TBA
Aug. 25, 2021
:
4 p.m. in Zoom
Abstract
We will hold an organizational meeting to welcome everyone in the first week. At this stage, the meeting is going to be a remote format on Zoom. The seminar access link and other information will be sent out through the seminar email list when the date comes closer.
Sept. 8, 2021
Danielle Tucker :
4 p.m. in Zoom
Abstract
Global Fréchet regression is an extension of linear regression to cover more general types of responses, such as distributions, networks and manifolds, which are becoming more prevalent. In such models, predictors are Euclidean while responses are metric space valued. Predictor selection is of major relevance for regression modeling in the presence of multiple predictors but has not yet been addressed for Fréchet regression. Due to the metric space valued nature of the responses, Fréchet regression models do not feature model parameters, and this lack of parameters makes it a major challenge to extend existing variable selection methods for linear regression to global Fréchet regression. In this work, we address this challenge and propose a novel variable selection method that overcomes it and has good practical performance. We provide theoretical support and demonstrate that the proposed variable selection method achieves selection consistency. We also explore the finite sample performance of the proposed method with numerical examples and data illustrations.
Sept. 15, 2021
Ted Westling :
4 p.m. in Zoom
Abstract
Much of the literature on estimating causal effects concerns discrete exposures. Recently, there has been increased interest in continuous exposures; that is, exposures that can take an uncountable number of values. Examples of such exposures include air pollution, pre-vaccination antibody responses, and concentrations of harmful chemicals in the blood. In this talk, I will provide an introduction to the area of causal inference with continuous exposures. I will then provide an overview of some of the recent research concerning nonparametric causal inference with continuous exposures, including my own recent and ongoing research. In particular, I will discuss approaches to nonparametric pointwise and global inference on causal dose-response curves, and, time permitting, inference on alternative causal parameters such as the effects of stochastic and incremental interventions.
Sept. 29, 2021
Xiongtao Dai :
4 p.m. in Zoom
Abstract
Functional data analysis concerns a sample of random functions, such as a collection of body growth trajectories. Dimension reduction tools, such as functional principal component analysis, are available to reduce and represent the infinite-dimensional functions. In this work, we are interested in estimating densities as functions, where each density comes from a subpopulation. For example, in the context of epidemiology, the age distributions of patients with different diseases is of central interest, where the disease defines a subpopulation. A key challenge comes from the highly variable sample sizes for different conditions, making the estimation of age profiles difficult for rare conditions. We propose a fully data-driven approach to estimate the densities without the need of specifying the parametric form of the density families. The idea is to map the density functions to a Hilbert space and then apply functional data analytic methods so as to derive low-dimensional approximates. I will show that the proposed methods yield interpretable results and are efficient for modeling electronic medical records and extreme rainfall.
Oct. 6, 2021
Hsin-Hsiung Huang :
4 p.m. in Zoom
Abstract
While matrix variate regression models have been studied in many existing works, classical statistical and computational methods for analysis of the regression coefficient estimation are highly affected by ultrahigh dimensional matrix-valued predictors. To address this issue, this paper proposes a framework of matrix variate regression methods, based on a rank-constraint optimization problem and its alternating gradient descent algorithm. In particular, we consider three low-rank matrix variate regression models including ordinary matrix regression, robust matrix regression, and matrix logistic regression, and we establish the convergence property and statistical consistency of the proposed estimator under these three models. The rank constraint effectively reduces the number of parameters in the model, and as a result, compared with existing methods based on regularization, our method has a better theoretical consistency rate. The experimental results show that the proposed algorithms are effective and efficient under various settings.
Oct. 13, 2021
Jonathan Niles-Weed :
4 p.m. in Zoom
Abstract
Given two probability distributions in R^d, a transport map is a function which maps samples from one distribution into samples from the other. For absolutely continuous measures, Brenier proved a remarkable theorem identifying a unique canonical transport map, which is monotone in a suitable sense. We study the question of whether this map can be efficiently estimated from samples. The minimax rates for this problem were recently established by Hutter and Rigollet (2021), but the estimator they propose is computationally infeasible in dimensions greater than three. We propose two new estimators---one minimax optimal, one not---which are significantly more practical to compute and implement. The analysis of these estimators is based on new stability results for the optimal transport problem and its regularized variants.
Based on joint work with Manole, Balakrishnan, & Wasserman and with Pooladian.
Oct. 20, 2021
Yanxi Liu :
4 p.m. in Zoom
Abstract
As the data size increases rapidly, the relationship between input and output variables may not be homogeneous anymore. Conventional statistical models such as generalized linear models (GLMs) may not be well-suited to heterogeneous relationships. Using a Mixture of Expert models is a good solution. The Mixture of Expert models can combine different statistical models to detect heterogeneous patterns while maintaining the benefits of conventional statistical modeling techniques. However, it needs a considerable amount of computer resources, particularly when working with big data. To address this issue, an attractive idea is to analyze a subsample of the data retaining the rich information of the full data. Information-Based Optimal Subdata Strategy (IBOSS), proposed by Wang et al. (2019), is such a strategy. The IBOSS strategy captures most of the relevant information in the full data through a judicious selection of the subdata by "maximizing" the Fisher information matrix. This project aims to develop an algorithm for the Clusterwise Linear Regression model, a type of Mixture of Experts, to select subdata based on IBOSS strategy. However, the Fisher information matrix of the model has no explicit form, which is a major challenge of the work. To overcome this challenge, we propose a surrogate matrix which is proved to be asymptotically equivalent to the Fisher information matrix, and it is used to construct the IBOSS subdata. Further, the proposed subdata selection is proved to be asymptotically optimal, i.e., no other method is statistically more efficient than the proposed one when the full data size is large.
Oct. 27, 2021
Youjin Lee :
4 p.m. in Zoom
Abstract
Instrumental variables have been widely used to estimate the causal effect of a treatment on an outcome in the presence of unmeasured confounders. When several instrumental variables are available and the instruments are subject to possible biases that do not completely overlap, a careful analysis based on these several instruments can produce orthogonal pieces of evidence (i.e., evidence factors) that would strengthen causal conclusions when combined. We develop several strategies, including stratification, to construct evidence factors from multiple candidate instrumental variables when invalid instruments may be present. Our proposed methods deliver nearly independent inferential results each from candidate instruments under the more liberally defined exclusion restriction than the previously proposed reinforced design. We apply our stratification method to evaluate the causal effect of malaria on stunting among children in Western Kenya using three nested instruments that are converted from a single ordinal variable. Our proposed stratification method is particularly useful when we have an ordinal instrument of which validity depends on different values of the instrument.
This is based on joint work with Anqi Zhao, Dylan Small, and Bikram Karmarkar.
Nov. 3, 2021
Bei Jiang :
4 p.m. in Zoom
Abstract
There is a growing expectation that data collected by government-funded studies should be openly available to ensure research reproducibility, which also increases concerns about data privacy. A strategy to protect individuals' identity is to release multiply imputed (MI) synthetic datasets with masked sensitivity values (Rubin, 1993). However, information loss or incorrectly specified imputation models can weaken or invalidate the inferences obtained from the MI-datasets. We propose a new masking framework with a data-augmentation (DA) component and a tuning mechanism that balances protecting identity disclosure against preserving data utility. Applying it to a restricted-use Canadian Scleroderma Research Group (CSRG) dataset, we found that this DA-MI strategy achieved a 0% identity disclosure risk and preserved all inferential conclusions. It yielded 95% confidence intervals (CIs) that had overlaps of 98.5% (95.5%) on average with the CIs constructed using the full, unmasked CSRG dataset in a work-disability (interstitial lung disease) study. The CI-overlaps were lower for several other methods considered, ranging from 73.9% to 91.9% on average with the lowest value being 28.1%; such low CI-overlaps further led to some incorrect inferential conclusions. These findings indicate that the DA-MI masking framework facilitates sharing of useful research data while protecting participants' identities.
This a joint work with Adrian Raftery (University of Washington), Russel Steele (McGill University) and Naisyin Wang (University of Michigan).
Nov. 10, 2021
Pang Du :
4 p.m. in Zoom
Abstract
We consider the problem of comparing probability densities between two groups. A new probabilistic tensor product smoothing spline framework is developed to model the joint density of two variables. Under such a framework, the probability density comparison is equivalent to testing the presence/absence of interactions. We propose a penalized likelihood ratio test for such interaction testing and show that the test statistic is asymptotically chi-square distributed under the null hypothesis. Furthermore, we derive a sharp minimax testing rate based on the Bernstein width for nonparametric two-sample tests and show that our proposed test statistics is minimax optimal. In addition, a data-adaptive tuning criterion is developed to choose the penalty parameter. Simulations and real applications demonstrate that the proposed test outperforms the conventional approaches under various scenarios.
Nov. 17, 2021
Yuefeng Han :
4 p.m. in Zoom
Abstract
Motivated by modern scientific research, analysis of tensors (multi-dimensional arrays) has emerged as one of the most important and active areas in modern statistics and data science. High-dimensional tensor data routinely arise in a wide range of applications, such as economics, genetics, microbiome studies, brain imaging, and hyperspectral imaging, due to modern data collection capabilities. In many of these settings, the observed tensors are of high dimension and high order, but the important information may lie in dimension-reduced subspaces induced by various structural conditions. This talk aims to develop new methodologies and theories from a perspective of subspace learning.
The talk is divided into two parts. In the first part, we introduce a factor approach for analyzing high dimensional dynamic tensors, in a form similar to Tucker tensor decomposition. We propose two estimation methods that are based on the tensor unfolding of lagged cross-product and iterative orthogonal projections of the original dynamic tensors. We also establish computational and statistical guarantees of the proposed methods. In the second part, we investigate a tensor factor model with a CP type low-rank tensor structure. We develop a new computationally efficient estimation procedure, which includes a warm-start initialization and an iterative concurrent orthogonalization scheme. We show that the iterative algorithm achieves $\epsilon$-accuracy guarantee within $\log\log(1/\epsilon)$ number of iterations.
Dec. 1, 2021
Cong Ma :
4 p.m. in Zoom
Abstract
This talk is concerned with the problem of off-policy evaluation in the multi-armed bandit model with bounded rewards. We develop minimax rate-optimal procedures under three different settings. First, when the behavior policy is known, we show that the Switch estimator, a method that alternates between the plug-in and importance sampling estimators, is minimax rate-optimal for all sample sizes. Second, when the behavior policy is unknown, we analyze performance in terms of the competitive ratio, thereby revealing a fundamental gap between the settings of known and unknown behavior policies. When the behavior policy is unknown, any estimator must have mean-squared error larger---relative to the oracle estimator equipped with the knowledge of the behavior policy---by a multiplicative factor proportional to the support size of the target policy. Moreover, we demonstrate that the plug-in approach achieves this worst-case competitive ratio up to a logarithmic factor. Third, we initiate the study of the partial knowledge setting in which it is assumed that the minimum probability taken by the behavior policy is known. We show that the plug-in estimator is optimal for relatively large values of the minimum probability, but is sub-optimal when the minimum probability is low. In order to remedy this gap, we propose a new estimator based on approximation by Chebyshev polynomials that provably achieves the optimal estimation error. This is a joint work with Banghua Zhu, Jiantao Jiao and Martin Wainwright.
Feb. 2, 2022
TBA :
4 p.m. in Zoom
Feb. 16, 2022
Liangbing Luo :
4 p.m. in Zoom
Abstract
In this talk, I will discuss logarithmic Sobolev inequalities with respect to a heat kernel measure on finite-dimensional and infinite-dimensional Heisenberg groups. Such a group is the simplest non-trivial example of a sub-Riemannian manifold. First, I will talk about logarithmic Sobolev inequalities on non-isotropic Heisenberg groups and discuss the dimension (in)dependence of the constants. In this setting, a natural Laplacian is not an elliptic but a hypoelliptic operator. The argument relies on comparing logarithmic Sobolev constants for the three-dimensional non-isotropic and isotropic Heisenberg groups, and tensorization of logarithmic Sobolev inequalities in the sub-Riemannian setting. Moreover, I will mention the application of these results to an infinite-dimensional Heisenberg group.
March 2, 2022
Yumou Qiu :
4 p.m. in Zoom
Abstract
We study statistical inference for the effects of multiple covariates of interest simultaneously after adjusting the effects of high-dimensional control variables under a linear model. A residual refitting procedure is proposed which first obtains the residuals from fitting the response variable and the target covariates on the control covariates via regularized estimation, and then refit the residuals from the first step. Hypothesis testing and confidence interval are constructed. The proposed procedure reduces the impact of the potential over-fitting errors from regularized estimation on the inference of the target parameters. It eliminates the prediction errors in the direction of the true regression error, and hence, achieving more accurate size and higher power. Expansions of the proposed statistics are derived without a sparsity condition on the precision matrix of covariates, which show the error reduction property of the residual refitting procedure. Simulation studies and real data analysis for S&P 500 stock returns verify the theoretical results and demonstrate the proposed method has better performance than the existing methods.
March 9, 2022
Rishi Sonthalia :
4 p.m. in Zoom
Abstract
Many important machine learning problems can be formulated as highly constrained convex optimization problems. One important example is metric constrained problems. In this paper, we show that standard optimization techniques can not be used to solve metric constrained problem.
To solve such problems, we provide a general active set framework, called Project and Forget, and several variants thereof that use Bregman projections. Project and Forget is a general purpose method that can be used to solve highly constrained convex problems with many (possibly exponentially) constraints. We provide a theoretical analysis of Project and Forge} and prove that our algorithms converge to the global optimal solution and have a linear rate of convergence.
In this talk I will go over the main details of the algorithm, the convergence results and applications to metric constrained and non metric constrained problems. For the non metric constrained problem I will present an application to a new formulation of unbalanced optimal transport known as dual regularized optimal transport.
March 30, 2022
Peijun Sang :
4 p.m. in Zoom
Abstract
We propose inferential tools for functional linear quantile regression where the conditional quantile of a scalar response is assumed to be a linear functional of a functional covariate. In contrast to conventional approaches, we employ kernel convolution to smooth the original loss function. The coefficient function is estimated under a reproducing kernel Hilbert space framework. A gradient descent algorithm is designed to minimize the smoothed loss function with a roughness penalty. With the aid of the Banach fixed-point theorem, we show the existence and uniqueness of our proposed estimator as the minimizer of the regularized loss function in an appropriate Hilbert space. Furthermore, we establish the convergence rate as well as the weak convergence of our estimator. As far as we know, this is the first weak convergence result for a functional quantile regression model. Pointwise confidence intervals and a simultaneous confidence band for the true coefficient function are then developed based on these theoretical properties. Numerical studies including both simulations and a data application are conducted to investigate the performance of our estimator and inference tools in finite sample. This is a joint work with my collaborators Zuofeng Shang and Pang Du.
April 6, 2022
Roland Molontay :
4 p.m. in Zoom
April 13, 2022
Yong Zeng :
4 p.m. in Zoom
Abstract
We propose a general partially-observed framework of Markov processes with marked point process observations for ultrahigh frequency (UHF) transaction price data, allowing other observable economic or market factors. We develop the corresponding Bayesian inference via filtering equations to quantify parameter and model uncertainty. Specifically, we derive filtering equations, which are SPDEs, to characterize the evolution of the statistical foundation such as likelihoods, posteriors, Bayes factors, and posterior model probabilities. Given the computational challenge, we provide a weak convergence theorem, enabling us to employ the Markov chain approximation method to construct consistent, easily-parallelizable, recursive algorithms. The algorithms calculate the fundamental statistical characteristics and are capable of implementing the Bayesian inference in real-time for streaming UHF data via parallel computing for sophisticated models. The general theory is illustrated by specific models built for U.S. Treasury Notes transactions data from GovPX and a Heston stochastic volatility model for stock transactions data. This talk consists of joint works with B. Bundick, G. X. Hu, D. Kuipers, and J. Yin.
April 20, 2022
Alexander Gutfraind :
4 p.m. in Zoom
Abstract
Randomized clinical trials are a pillar of evidence-based medicine, but are often very expensive and recruit a poor representation of the at-risk population. Here we seek to develop a recruitment strategy for trials of vaccines for HCV that would not only decrease the required sample size to achieve adequate statistical power, but also improve the demographic representation of the recruited trial cohort to enhance their equity and generalizability. Using PWID data collected from Chicago, predictive incidence models were trained and applied to a recruitment scheme which aggregates and enrolls candidates in a batchwise manner and incorporates sample size re-estimation. Dynamic weights are applied to generate a numerical score that can be used to assess a candidate’s expected probability of infection and demographic desirability, thus allowing trials to selectively recruit high-incidence PWID who also contribute to the generalizability of the trial. Simulated clinical trial recruitment using this scheme expressed a two- to three-fold increase in HCV incidence among the trial cohort compared to conventional methods. Simultaneously, the demographic composition of the recruited cohort more closely resembled the target population. This recruitment scheme also proved flexible to varying numbers of matched demographic categories, while also being robust to target populations that were highly dissimilar to the recruitment pool. This novel method of trial recruitment presents a promising approach by which costs can be minimized while assuring a high level of demographic representation. This is joint work with Richard Guan Chiu.
April 27, 2022
Jian Zhao :
4 p.m. in Zoom
Abstract
This presentation will describe the roles of statisticians at FDA, Center for Drug Evaluation and Research. High level overview of FDA missions and organizations as well as responsibilities of FDA statisticians in the Office of Biostatistics will be presented. IND and NDA review process will be briefly introduced followed by a review case study for illustration. In addition, a variety of opportunities for statisticians working at FDA including public meetings, guidance & policy development, working groups, research, leadership development, will be discussed.
Aug. 31, 2022
TBA :
4 p.m. in 636 SEO
Abstract
We will hold a welcome meeting for everyone. The new students are encouraged to join to get information about the seminar format.
Sept. 7, 2022
Qunfeng Dong :
4 p.m. in 636 SEO
Abstract
Many biomedical researchers including myself lack the formal training in Bayesian statistics, yet we have found the beauty and magic in it. In this talk, I will present three projects: (1) Bayesian modeling to estimate hospitalization risk for COVID-19 patients with comorbidities, (2) a microbiome taxonomic classification method based on Bayes theorem and bootstrapping, and (3) predicting clinical outcomes of metastatic melanoma patients based on the commensal microbiome using a Bayes’ classifier. I will also highlight the limitations of our methods, so that hopefully hardcore mathematicians/statisticians can come up with better solutions.
References:
1. Xiang Gao and Qunfeng Dong (2020) A Bayesian Framework for Estimating the Risk Ratio of Hospitalization for People with Comorbidity Infected by the SARS-CoV-2 Virus. Journal of the American Medical Informatics Association, 28 Sept 2020, ocaa246, doi:10.1093/jamia/ocaa246
2. Xiang Gao, Huaiying Lin, Qunfeng Dong (2017); A Dirichlet-Multinomial Bayes Classifier for Disease Diagnosis with Microbial Compositions, mSphere, Volume: 2, Issue: 6.
3. Xiang Gao, Huaiying Lin, Kashi Revanna, Qunfeng Dong (2017) A Bayesian Taxonomic Classification Method for 16S rRNA Gene Sequences with Improved Species-level Accuracy. BMC Bioinformatics 2017 May 10;18(1):247.
Sept. 21, 2022
Nabil Kahouadji :
4 p.m. in 636 SEO
Abstract
We introduce twenty-four new two-parameter families of advanced time series forecasting functions, using three forecast estimate methods along with eight optimization criteria. We also introduce the concept of powering and derive non-seasonal and seasonal time series models with examples in education, sales, economics, industry and finance. We compare the performance of our twenty-four functions/models to both exponential smoothing and ARIMA models using non-seasonal and seasonal time series. We show in particular that our models not only do not require a decomposition of a seasonal time series into trend, seasonal and random components, but also leads to substantially lower sum of absolute error and a higher number of closer forecasts than both Holt--Winters and ARIMA models. Finally, we apply and compare the performance of our twenty-four models using five-year stock market data of 467 companies of the S&P500.
Oct. 5, 2022
Roberto Molinari :
4 p.m. in 636 SEO
Abstract
Differential privacy (DP) provides an elegant mathematical framework for defining a provable disclosure risk in the presence of arbitrary adversaries: it guarantees that whether an individual is in a database or not, the results of a DP procedure should be similar in terms of their probability distribution. While DP mechanisms are provably effective in protecting privacy, they often negatively impact the utility of the query responses, statistics, and/or analyses that come as outputs from these mechanisms. To address this problem, we use ideas from the area of robust statistics, which aims at reducing the influence of outlying observations on statistical inference. Based on the preliminary known links between differential privacy and robust statistics, we modify the objective perturbation mechanism by making use of a new bounded function and define a bounded M-Estimator with adequate statistical properties. The resulting privacy mechanism, named “Perturbed M-Estimation”, shows important potential in terms of improved statistical utility of its outputs as suggested by some preliminary results. These results motivate the current work which is being made in this direction.
Oct. 12, 2022
Renming Song :
4 p.m. in 636 SEO
Abstract
In this talk, I will present some recent results on potential theory of
Dirichlet forms on the half-space $\mathbb{R}^d_+$ defined by the jump kernel
$J(x,y)=|x-y|^{-d-\alpha}{\cal B}(x,y)$, where $\alpha\in (0,2)$ and
${\cal B}(x,y)$ can blow up to infinity at the boundary. The main results
include boundary Harnack principle and sharp two-sided Green function estimates.
This talk is based on a joint paper with Panki Kim and Zoran Vondracek.
Oct. 19, 2022
Roland Molontay :
4 p.m. in Zoom
Abstract
Anomaly detection refers to the process of identifying unexpected objects or patterns, which do not conform to the usual behavior. The detection of “not-normal” observations has attracted a lot of research interest from the machine learning community since it has a wide variety of practical applications.
In this talk, I will briefly present an overview of the challenges of unsupervised anomaly detection. I will also present our novel model-based approach that relies on the multivariate probability distribution associated with the observations [1]. Since the rare events are present in the tails of the probability distributions, we use copula functions, which are able to model the fat-tailed distributions well. The presented procedure scales well; it can cope with a large number of high-dimensional samples and also with missing values.
I will also demonstrate the usability of the method through a case study, where we analyze a large dataset consisting of the performance counters of a real mobile telecommunication network.
[1]: Horváth, G., Kovács, E., Molontay, R., & Nováczki, S. (2020). Copula-based anomaly scoring and localization for large-scale, high-dimensional continuous data. ACM Transactions on Intelligent Systems and Technology (TIST), 11(3), 1-26
Oct. 26, 2022
Yuehua Cui :
4 p.m. in 636 SEO
Yuehua Cui :
4 p.m. in Zoom
Abstract
Mendelian randomization (MR) uses genetic variants as instrument variables to determine whether an observational association between an exposure and an outcome is causal. The use of Mendelian Randomization reduces regression bias and provides reliable estimate of the likely underlying causal relationship between an exposure and a disease outcome. Most current Mendelian randomization methods are focused on cross-sectional phenotypic traits. Longitudinal studies track the same individual at different time points and have a number of advantages over cross-sectional studies. Motivated by a real study to evaluate the causal effect of hormone level on eating behavior, we propose two MR models to investigate the causal effects in a longitudinal study. In the first model, we assume the current exposure affects the current outcome. In the second model, we assume that the past and/or current exposures contribute to the current outcome. The delayed causal effect is determined by data through a variable selection algorithm. Point-wise and simultaneous testing are developed to assess the existence of causal effects. The method was illustrated via simulation studies and an application to an eating behavior dataset.
Nov. 2, 2022
Soo-Young Kim :
4 p.m. in Zoom
Abstract
Various methods for estimating variance functions in heteroscedastic regression models have been developed over the years. We propose methods to estimate a variance function in a heteroskedastic regression model where the variance function is assumed to be smooth and monotone in a predictor variable. The estimation method is based on the maximum likelihood principle, and its computation is carried out through regression splines and the cone projection algorithm. The convergence rate of the estimated variance function is derived, and simulations show that it tends to be closer to the true variance function in a variety of scenarios compared to the existing methods. The estimated variance function from the proposed method provides improved inference about the mean function, in terms of a coverage probability and an average length for an interval estimate. The utility of the method is illustrated through the analysis of real datasets.
Nov. 9, 2022
Hsin-Hsiung Huang :
4 p.m. in Zoom
Abstract
Inspired by our recent works on the NSF ATD challenges for spatiotemporal data analysis and modeling and Bayesian clustering research, we investigate whether the Bayesian methods can consistently estimate the model parameters when there are multivariate mixed-type responses. To this end, shrinkage priors are useful for identifying relevant signals in high-dimensional data. We develop a multivariate Bayesian model with shrinkage priors (MBSP) model to mixed-type response generalized linear models (MRGLMs), and we consider a latent multivariate linear regression model associated with the observable mixed-type response vector through its link function. Under our proposed model (MBSP-GLM), multiple responses belonging to the exponential family are simultaneously modeled and mixed-type responses are allowed. We show that the MBSP-GLM model achieves strong posterior consistency when $p$ grows at a subexponential rate with $n$. Furthermore, we quantify the posterior contraction rate at which the posterior shrinks around the true regression coefficients and allow the dimension of the responses $q$ to grow as $n$ grows. This greatly expands the scope of the MBSP model to include response variables of many data types, including binary and count data. To address the non-conjugacy concern, we propose an adaptive sampling algorithm via a P\'{o}lya-gamma data augmentation scheme for the MRGLM estimation. We provide simulation studies and real data examples.
Feb. 22, 2023
Hani Aldirawi :
4:30 p.m. in 636 SEO
Abstract
Sparse data with a high portion of zeros arise in various disciplines. Modeling sparse data is a challenging and growing research area. In this presentation, we provide statistical methods and tools for analyzing sparse discrete data such as microbiome data. We aim to answer three main questions related to microbiome data. 1) What is the most appropriate probabilistic model for modeling each single microbiome feature? 2) How do we build the most appropriate regression model for each feature when covariates are available? 3) Suppose we have repeated measurements for each subject (longitudinal data), how can we identify the time intervals when the two groups of individuals are significantly different?
March 15, 2023
Ming-Chung Chang :
6 p.m. in Zoom
Abstract
Predictive analytics encompasses the use of statistical models for prediction. Its power, however, is hindered by the rising amounts of data in recent years. Owing to advanced technology, big data are ubiquitous across disciplines. Such data richness may yield difficulties in predictive analytics either in terms of time cost or numerical stability. In this talk, I will introduce a new subsampling approach to overcome this difficulty for regression problems. The proposed method integrates a nonparametric regression technique and stratified sampling, referred to as supervised stratified subsampling. Theoretical properties are developed to justify this method. Numerical studies show that the proposed method yields good predictions and is against model misspecification.
April 5, 2023
Kai Zhang :
4 p.m. in 636 SEO
Abstract
We study the problem of distribution-free dependence detection and modeling through the new framework of binary expansion statistics (BEStat). The binary expansion testing (BET) avoids the problem of non-uniform consistency and improves upon a wide class of commonly used methods (a) by achieving the minimax rate in sample size requirement for reliable power and (b) by providing clear interpretations of global relationships upon rejection of independence. The binary expansion approach also connects the symmetry statistics with the current computing system to facilitate efficient bitwise implementation. Modeling with the binary expansion linear effect (BELIEF) is motivated by the fact that wo linearly uncorrelated binary variables must be also independent. Inferences from BELIEF are easily interpretable because they describe the association of binary variables in the language of linear models, yielding convenient theoretical insight and striking parallels with the Gaussian world. With BELIEF, one may study generalized linear models (GLM) through transparent linear models, providing insight into how modeling is affected by the choice of link. We explore these phenomena and provide a host of related theoretical results. This is joint work with Benjamin Brown and Xiao-Li Meng.
April 12, 2023
Boxiang Wang :
4 p.m. in 636 SEO
Abstract
Cluster analysis is a fundamental task in machine learning. Several clustering algorithms have been extended to handle high-dimensional data by incorporating a sparsity constraint in the estimation of a mixture of Gaussian models. Though it makes some neat theoretical analysis possible, this type of approach is arguably restrictive for many applications. In this talk, I will introduce a novel latent variable transformation mixture model for clustering in which a mixture of Gaussians is assumed after some unknown monotone data transformation. A new clustering algorithm named CESME is developed for high-dimensional clustering under the assumption that optimal clustering admits a sparsity structure. The use of unspecified transformation makes the model far more flexible than the classical mixture of Gaussians. On the other hand, the transformation also brings quite a few technical challenges to the model estimation as well as the theoretical analysis of CESME. I will present a comprehensive analysis of CESME including identifiability, initialization, algorithmic convergence, and statistical guarantees on clustering. In addition, the convergence analysis has revealed an interesting algorithmic phase transition for CESME, which has also been noted for the EM algorithm in the literature. Leveraging such a transition, a data-adaptive procedure is developed and substantially improves the computational efficiency of CESME. Extensive numerical study and real data analysis show that CESME outperforms the existing high-dimensional clustering algorithms including CHIME, sparse spectral clustering, sparse K-means, sparse convex clustering, and IF-PCA.
April 19, 2023
Yimin Xiao :
4 p.m. in 636 SEO
Abstract
Local times of a Gaussian random field $X = \{X(t),t ∈ \mathbb{R}^N\}$ with values in $\mathbb{R}^d$ carry a lot of analytic and geometric properties about $X$. They also arise naturally in the limit distributions of functionals of integrated and fractionally integrated time series or spatial processes, and in nonlinear cointegrating regression.
In this talk, we study the local times of anisotropic Gaussian random fields satisfying strong local nondeterminism with respect to an anisotropic metric. By applying moment estimates for local times, we prove optimal local and global Hölder conditions for the local times for these Gaussian random fields and deduce related sample path properties. These results are closely related to Chung’s law of the iterated logarithm and the modulus of nondifferentiability of the Gaussian random fields.
We apply the results to systems of stochastic heat equations with additive Gaussian noise and determine the exact Hausdorff measure function for the level sets of the solution.
This talk is based on a joint paper with Davar Khoshnevisan and Cheuk Yin Lee.
April 26, 2023
Xiang Zhu :
4 p.m. in 636 SEO
Abstract
Large-scale genome-wide association studies (GWAS) have markedly improved our understanding of how common variation in the human genome affects complex traits and diseases. Regression models have been widely used to analyze GWAS, but existing methods often require input data at the individual level, which are hard to obtain due to many administrative issues. Here we provide a Bayesian framework for multiple regression without the need of individual-level data. Specifically, we derive a "Regression with Summary Statistics" (RSS) likelihood function of the multiple regression coefficients based on the univariate regression summary statistics, which are easily available in GWAS. We combine the RSS likelihood with prior distributions that are specifically designed for a wide range of genetic applications, such as heritability estimation, phenotype prediction, pathway enrichment and gene prioritization. To estimate posterior distributions, we develop efficient Markov chain Monte Carlo and variational inference algorithms that scales well with millions of genetic variants. Applying RSS to a host of real-world GWAS summary statistics, we demonstrate that RSS not only achieves similar performance in settings where existing methods work, but also enables many novel analyses and discoveries that existing methods cannot deliver.
Aug. 30, 2023
TBA :
4 p.m. in 636 SEO
Abstract
In this welcome meeting, the modality of the seminar will be discussed in person. Everyone is invited. The new students are encouraged to attend to get information and communicate with future colleagues.
Sept. 13, 2023
Sergey Tarima :
4 p.m. in 636 SEO
Abstract
The possibility of early stopping and/or interim sample size re-estimation lead to random sample sizes. When such interim adaptations are informative, the interim decision becomes a component of the sufficient statistic. We decompose the total Fisher Information (FI) into the design FI and a conditional-on-design FI analogous to Molenberghs et al. (2014). We go further, representing the conditional-on-design FI as a weighted linear combination of FIs conditional on realized decisions. This decomposition is useful for quantifying how much mean-squared error will be lost due to planned-informative adaptations. We use The FI unspent by having a planned-informative adaptation to determine the lower bound on mean squared error of post-adaptation estimators [the Cramer-Rao lower bound (1946) and its sequential version suggested by Wolfowitz (1947) are not applicable to such estimators]. Theoretical results are illustrated with simple normal samples collected according to a two-stage design with a possibility of early stopping.
Sept. 27, 2023
T.E.S. Raghavan :
4 p.m. in 636 SEO
Abstract
Often graduate students concentrate on accumulating high grades in a variety of graduate courses and years pass by with no real thesis problem in sight. ISI style has always been to let you struggle till you discover your own thesis problem. Inspired by the first chapter of Wald’s seminal work on Statistical decision functions, when I jumped to chapter 2, I found my initial mathematical background was quite inadequate. in steering my initial interest from Wald’s Statistical decision theory into Game theory and to the theory of positive operators, many great researchers at ISI Kolkatta have played a significant role.
My talk will walk through some facets of this struggle and the decisive role played by professors C.R Rao, D. Basu, K.R Parthasarathy, V.S Varadarajan and S.R. S Varadhan.
Oct. 18, 2023
Junhyeon Kwon :
4 p.m. in Zoom
Abstract
Hawkes process is one of the most commonly used models for investigating the self-exciting nature of earthquake occurrences. However, seismicity patterns have complicated characteristics due to heterogeneous geology and stresses, for which existing methods with Hawkes process cannot fully capture. This study introduces novel nonparametric Hawkes process models that are flexible in three distinct ways. First, we incorporate the spatial inhomogeneity of the self-excitation earthquake productivity. Second, we consider the anisotropy in aftershock occurrences. Third, we reflect the space–time interactions between aftershocks with a non-separable spatio-temporal triggering structure. For model estimation, we extend the model-independent stochastic declustering (MISD) algorithm and suggest substituting its histogram-based estimators with kernel methods. We demonstrate the utility of the proposed methods by applying them to the seismicity data in regions with active seismic activities.
Nov. 15, 2023
Lingjie Ma :
4 p.m. in 636 SEO
Abstract
This paper focuses on portfolio construction at tails. The classical mean-variance portfolio focuses on the first two moments of a return distribution.
However, asset returns usually are not Gaussian distributed; rather, they have long and fat tails.
As a complementary approach, this paper studies and incorporates tail risk into portfolio optimization. Using 1970 to 2019 S&P 500 data, an empirical study was performed by constructing realistic stock selection investment strategies. The results indicate that quantile optimization produces practical tail portfolios with volatility, diversity and turnover comparable to the classical mean-variance approach. Moreover, a tail portfolio outperforms the mean-variance portfolio consistently over the period studied, particularly when the market is bearish.
Nov. 22, 2023
Dr. Xianwei Bu :
4 p.m. in 636 SEO
Abstract
This oral presentation introduces the clinical trials and new drug development in a pharmaceutical company and a statistician’s roles and responsibilities during the process. Some commonly used statistical analysis methods in clinical trials are described. Priorities and timelines are highlighted as two features for a statistician, followed by a summary of how to become an effective statistician in a pharmaceutical company.
Jan. 31, 2024
Wenxin Zhou :
4 p.m. in 636 SEO
Abstract
Expected Shortfall (ES), also known as superquantile or Conditional Value-at-Risk, has been recognized as an important measure in risk analysis and stochastic optimization. In finance, it refers to the conditional expected return of an asset given that the return is below some quantile of its distribution. In this talk, we consider a joint regression framework that simultaneously models the conditional quantile and ES of a response variable given a set of covariates, for which the state-of-the-art approach is based on minimizing a joint loss function that is non-differentiable and non-convex.
Motivated by the idea of using Neyman-orthogonal scores to reduce sensitivity with respect to nuisance parameters, we propose statistically robust and computationally efficient two-step procedures for fitting joint quantile and ES regression models under three settings: (i) the classical linear model with $p\ll n$; (ii) high-dimensional sparse models with $p\gg n$, and (iii) nonparametric models with a hierarchical compositional structure. Furthermore, we discuss a more general integrated-quantile regression framework, including ES regression as a special case.
Feb. 14, 2024
Moontae Lee :
4 p.m. in Zoom
Abstract
Large Language Models (LLMs) have transformed Natural Language Processing and the wider spectrum of Artificial Intelligence. Relying on their capability to understand extensive contexts, the groundbreaking innovation lies in tackling multiple tasks by a single formalism: generating contextually coherent and creatively diverse subsequent outputs. This talk overviews recent progress in large language modeling and my journey into the realm of generative AI. Highlighting my focuses on both text and code modalities, the talk delves into my recent research on retrieval-augmented text generation and multilingual code completions. Furthermore, the presentation outlines state-of-the-art ongoing research trajectory on planning, reasoning, aligning, and prospective steps toward self-learning. In conclusion, the talk also touches critical problems in Safety, Ethics, and Governance in Artificial Intelligence.
Feb. 21, 2024
Hyo Young Choi :
4 p.m. in Zoom
Abstract
Over the last decade, many innovative technologies have generated vast amounts of large-scale biological data. The accumulation of so-called “big data”, especially from next generation sequencing technologies, has created many exciting areas in statistics as well as biology. In particular, statistical tools and machine learning techniques have proven to be critical in cancer genomics, transforming large and complex data into clinically relevant knowledge. While many computational tools have been developed for analyzing such big data, unprecedented challenges remain in turning it into meaningful and actionable insights. This talk primarily concerns the issue of high-dimensional outliers which are often challenging to identify in high-throughput sequencing data due to the special structure of high dimensional space. We introduce a new notion of high dimensional outliers that embraces various types and provides deep insights into understanding the behavior of these outliers based on several asymptotic regimes. As an important application, we introduce a statistical method for unsupervised screening of a range of structural alterations in RNA-seq data. We identify a number of biologically important outliers along with the successful characterization of the subspace associated with outliers, which holds promise for identifying otherwise obscured signals.
March 6, 2024
Dehan Kong :
4 p.m. in 636 SEO
Abstract
Understanding causal relationships is one of the most important goals of modern science. So far, the causal inference literature has focused almost exclusively on outcomes coming from the Euclidean space. However, it is increasingly common that complex biomedical datasets are best summarized as data points in non-linear spaces. In this paper, we present a novel framework of causal effects for outcomes from the Wasserstein space of cumulative distribution functions, which in contrast to the Euclidean space, is non-linear. We develop doubly robust estimators and associated asymptotic theory for these causal effects. As an illustration, we use our framework to quantify the causal effect of marriage on physical activity patterns using wearable device data collected through the National Health and Nutrition Examination Survey.
March 13, 2024
Yongzhao Shao :
4 p.m. in 636 SEO
Abstract
Alzheimer’s disease (AD) stands as the leading cause of dementia and related death. Presently, there is no cure or an effective prevention. AD pathology is highly heterogeneous which may begin 20 years before clinical diagnoses. Effective blood-based biomarkers are desired for early diagnosis and monitoring. Epigenetic clocks and mitotic clocks are individual-level biomarkers of aging and often referred as biological ages. Advanced age is known as the most impactful risk factor for late-onset AD, thus, biological ages have the potential to be the blood-based biomarkers of AD for diagnosis and monitoring in personalized medicine. In this talk, we will discuss the derivation of the epigenetic clocks based on DNA methylation profiles and the accelerations of biological ages as well as the mitotic clocks based on lymphocyte telomere lengths (LTLs). We will discuss causal analyses of the relationships between the epigenetic clocks, mitotic clocks and the risk of AD using Mendelian randomization (MR) analysis. The MR-based analyses using selected variants from genome-wide association studies as instrument variables are largely free of biases due to numerous measured and unmeasured confounding factors. If time permits, we will also discuss some MR-based causal analyses of the utility of heart failure medications in reducing the risk of AD among heart failure survivors. This talk is based on joint research with Dr. Yibeltal Ashebir and Jiehui Xu at NYU Grossman School of Medicine.
March 27, 2024
Yiou Li :
4 p.m. in 636 SEO
Abstract
Experimental designs for a generalized linear model (GLM) often depend on the specification of the model, including the link function,
the predictors, and unknown parameters, such as the regression coefficients. To deal with the uncertainties of these model specifications,
it is important to construct optimal designs with high efficiency under such uncertainties. Existing methods such as Bayesian experimental designs often use prior distributions of model specifications to incorporate model uncertainties into the design criterion. Alternatively, one can obtain the design by optimizing the worst-case design efficiency with respect to the uncertainties of model specifications. In this work, we propose a new Maximin Φp-
Efficient (or Mm-Φp for short) design which aims at maximizing the minimum Φp-efficiency under model uncertainties. Based on
the theoretical properties of the proposed criterion, we develop an efficient algorithm with sound convergence properties to construct the Mm-Φp design. The performance of the proposed Mm-Φp design is assessed through several numerical examples.
April 3, 2024
Hsin-Hsiung Huang :
4 p.m. in 636 SEO
Abstract
High-dimensional data have become prevalent in all fields that need statistical modeling and data analysis. I introduce my recent research in Bayesian ultrahigh dimensional variable selection, low-rank matrix regression and classification, and robust sufficient dimension reduction (SDR). We develop a Bayesian framework for mixed-type multivariate regression with continuous shrinkage priors that enables joint analysis of mixed continuous and discrete outcomes, allowing variable selection from a large number of covariates (p). We investigate the conditions for posterior contraction, especially when the number of covariates (p) grows exponentially relative to the sample size (n) and develop a two-step approach for variable selection with theorems of a sure screening property and posterior contraction and applications with simulation studies and applications to real datasets.
To address challenges in analyzing regression coefficient estimation affected by high-dimensional matrix-valued covariates, we propose a framework for matrix-covariate regression and classification models with a low-rank constraint and additional regularization for structured signals, considering continuous and binary responses, introduce an efficient Riemannian-steepest-descent algorithm for regression coefficient estimation, and prove the consistency of the proposed estimator, showing improvement over existing work in cases where the rank is small with applications through simulations and real datasets of shape images, brain signals, and microscopic leucorrhea images. We propose a novel SDR method robust against outliers using α-distance covariance that effectively estimates the central subspace under mild conditions on predictors without estimating a link function, based on the projection on the Stiefel manifold. We establish convergence properties of the proposed estimation under certain regularity conditions and compare the method's performance with existing SDR methods through simulations and real data analysis, highlighting improved computational efficiency and effectiveness.
April 10, 2024
Thomas Mathew :
4 p.m. in 636 SEO
Abstract
Identifying treatments or interventions that are cost-effective (more effective at a reasonable cost) is clearly important in health policy decision making, especially in the allocation of health care resources. Various measures of cost-effectiveness that are informative, intuitive and simple to explain have been suggested in the literature. Popular and widely used measures include the incremental cost-effectiveness ratio (ICER), defined as the ratio between the difference of average costs and the difference of average effectiveness in two populations receiving two treatments. The ICER is interpreted as the additional cost per unit of effectiveness gained. Yet another measure is the incremental net benefit (INB), which is the difference between the incremental cost and the incremental effectiveness after multiplying the latter with a "willingness-to-pay" amount. In the talk, I will provide a selected review of the statistical criteria and methodologies for cost-effectiveness analysis. In particular, some of the recently introduced probabilistic criteria will be discussed and examples will be given.
April 17, 2024
Lin Wang :
4 p.m. in 636 SEO
Abstract
Despite the availability of extensive data sets, it is often impractical to observe the responses or labels for all data points due to various measurement constraints in many applications. To address this challenge, subsampling approaches can be employed to select a subset of design points from a large pool for observation, resulting in substantial savings in labeling costs. In this presentation, I will introduce our recent research on computationally feasible subsampling techniques. Our primary focus is on regression with labeled data, which includes linear regression, ridge regression, and nonparametric additive regression. For these regression tasks, we have developed sampling probabilities that aim to minimize the mean squared error in estimations and predictions. We will demonstrate the effectiveness of our proposed approaches through both theoretical analysis and extensive simulations.
April 24, 2024
Kyunghee Han :
4 p.m. in 636 SEO
Abstract
The Statistical Laboratory will present the 2023-2024 Statistics Graduate Student Service Award, Consulting Award, and Research Award.
Sept. 4, 2024
TBA :
4 p.m. in 636 SEO
Abstract
In this meeting, the modality of the seminar will be discussed in person. Everyone is invited. The new students are encouraged to attend to get information and communicate with future colleagues.
Sept. 11, 2024
Dr. Hong Li :
4 p.m. in 636 SEO
Abstract
Bringing historical control information into a new trial appropriately holds the promise of more efficient trial design with more accurate estimates, increased power, and fewer patients allocated to inefficacious control group, provided the historical control data are sufficiently similar to the concurrent control. Interest has been growing over the past few decades in leveraging historical clinical trial on the control arm. However, most of the current historical borrowing methods focus on incorporating patient-level historical control information at only one time point. In this work, we propose a Bayesian hierarchical Mixed effect Models for Repeated Measures (BMMRM) to incorporate aggregated study-level longitudinal historical control estimates into the concurrent trial that collected repeated longitudinal data. The simulation study demonstrates that, as compared to one time point data analysis approach, leveraging longitudinal historical control data produces greater power enhancement and mitigates the power loss when the missing data under missing at random (MAR) mechanism is present. Our work also helps fill the gap of lack of methods borrowing historical longitudinal control data from the published summarized estimates when patient-level control data are not available.
Sept. 18, 2024
Seyoon Ko :
4 p.m. in Zoom
Abstract
Hawkes stochastic point process models have emerged as valuable statistical tools for analyzing viral contagion. The spatiotemporal Hawkes process characterizes the speeds at which viruses spread within human populations. Unfortunately, likelihood-based inference using these models requires O(N^2) floating-point operations, for N the number of observed cases. Recent work responds to the Hawkes likelihood's computational burden by developing efficient graphics processing unit (GPU)-based routines that enable Bayesian analysis of tens-of-thousands of observations. We build on this work and develop a high-performance computing (HPC) strategy that divides 30 Markov chains between 4 GPU nodes, each of which uses multiple GPUs to accelerate its chain's likelihood computations. We use this framework to apply two spatiotemporal Hawkes models to the analysis of one million COVID-19 cases in the United States between March 2020 and June 2023. In addition to brute-force HPC, we advocate for two simple strategies as scalable alternatives to successful approaches proposed for small data settings. First, we use known county-specific population densities to build a spatially varying triggering kernel in a manner that avoids computationally costly nearest neighbors search. Second, we use a cut-posterior inference routine that accounts for infections' spatial location uncertainty by iteratively sampling latent locations uniformly within their respective counties of occurrence, thereby avoiding full-blown latent variable inference for 1,000,000 infection locations.
Oct. 2, 2024
Bradley Jones :
4 p.m. in 636 SEO
Abstract
Definitive Screening Designs (DSDs) were introduced in 2011. Since then they have become popular for industrial applications. This talk describes what a DSD is. It then explains why engineers prefer them to standard two-level fractional factorial designs. Finally, it shows how to construct them and block them.
Oct. 9, 2024
Hyebin Song :
4 p.m. in 636 SEO
Abstract
In this talk, I will introduce a novel weighted l2 projection method for estimating covariance functions, with an emphasis on estimation of autocovariance sequences from reversible Markov chains. Shape-constrained estimation of a function with discrete support has been investigated and successfully applied to various application problems. Notably, Berg and Song (2023) connected this idea with uncertainty quantification in Markov chain Monte Carlo (MCMC) samples and proposed a shape-constrained estimator for autocovariance sequences. While the least-squares objective is commonly used in shape-constrained regression, it can be suboptimal due to correlation and unequal variances in the input function. To address this, we introduce a weighted least-squares method that defines a weighted norm on transformed data. Our approach involves transforming input data into the frequency domain and weighting the input sequence based on their asymptotic variances, exploiting the asymptotic independence of periodogram ordinates. I will discuss the computational aspects, theoretical properties, and the improved performance of this method compared to its non-weighted counterpart.
Oct. 16, 2024
Paromita Dubey :
4 p.m. in 636 SEO
Abstract
We introduce a powerful scan statistic and the corresponding test for detecting the presence and pinpointing the location of a change point within the distribution of a data sequence with the data elements residing in a separable metric space (Ω, d). These change points mark abrupt shifts in the distribution of the data sequence as characterized using distance profiles, where the distance profile of an element ω ∈ Ω is the distribution of distances from ω as dictated by the data. This approach is tuning parameter free, fully non-parametric and universally applicable to diverse data types, including distributional and network data, as long as distances between the data objects are available. We obtain an explicit characterization of the asymptotic distribution of the test statistic under the null hypothesis of no change points, rigorous guarantees on the consistency of the test in the presence of change points under fixed and local alternatives and near-optimal convergence of the estimated change point location, all under practicable settings. To compare with state-of-the-art methods we conduct simulations covering multivariate data, bivariate distributional data and sequences of graph Laplacians, and illustrate our method on real data sequences of the U.S. electricity generation compositions and Bluetooth proximity networks.
Oct. 23, 2024
Guanglei Hong :
4 p.m. in 636 SEO
Abstract
In education, health, and human services, an intervention program is usually implemented by many local organizations. Determining which organizations are more effective is essential for theoretically characterizing effective practices and for intervening to enhance the capacity of ineffective organizations. In multisite randomized trials, site-specific intention-to-treat (ITT) effects are likely invalid indicators for organizational effectiveness and may lead to inequitable decisions. This is because sites differ in their local ecological conditions including client composition, alternative programs, and community context. Applying the potential outcomes framework, this study proposes a mathematical definition for the relative effectiveness of an organization. The estimand contrasts the performance of a focal organization with those that share the features of its local ecological conditions. The identification relies on relatively weak assumptions by leveraging observed control group outcomes that capture the confounding impacts of alternative programs and community context. We propose a two-step mixed-effects modeling (2SME) procedure. Simulations demonstrate significant improvements when compared with site-specific ITT analyses or analyses that only adjust for between-site differences in the observed baseline participant composition. We illustrate its use through an evaluation of the relative effectiveness of individual Job Corps centers by reanalyzing data from the National Job Corps Study, a multisite randomized trial that included 100 Job Corps centers nationwide serving disadvantaged youths. The new strategy promises to alleviate consequential misclassifications of some of the most effective Job Corps centers as least effective and vice versa.
Nov. 6, 2024
Dr. Mandy Jin :
4 p.m. in Zoom
Abstract
In oncology clinical trials, subjects prematurely discontinuing from the assigned treatment prior to experiencing an event of interest are often handled by noninformative censoring under censor-at random assumption. Such methods can be challenged with respect to the robustness of the ignorable or noninformative censoring and sensitivity analyses using informative censoring are often required.
In a recently published article (Jin and Fang, 2024), reference-based methods (including Jump to Reference and Copy Reference) and tipping point analysis for time-to-event data with possibly informative censoring were proposed. These are novel methods to fit the gap in literature for time-to-event analysis with applications in oncology clinical trials. We will describe and facilitate the implementation of these methods in this presentation.
Illustrative examples are provided to demonstrate the reference-based methods and tipping point analysis.
Nov. 13, 2024
Dr. Yingda Lu :
4 p.m. in 636 SEO
Abstract
Consumers have increasingly spent more time on mobile applications, and companies have also allocated more resources to advertisement in mobile applications and are actively seeking ways to improve the click-through rate of in-app ads. However, there is a lack of research leveraging consumers’ mobile application usage to understand in-app advertisement. In this study, we develop an integrated model of mobile application usage and in-app advertising response. We use a hidden-Markov model (HMM), which allows consumer involvement in mobile activities to drive temporal changes in both consumer mobile application usage and in-app advertising response. Our framework captures three components that are understudied in previous research on in-app advertising responses: 1) contextual mobile app in which consumers are targeted; 2) long-range correlation in preceding periods and 3) multitasking across mobile apps. To address the challenge of long-range correlation in traditional HMM, we further extend HMM by incorporating a long short-term memory (LSTM) autoencoder into the state transition. Using a unique panel dataset, we find salient temporal patterns and persistence of consumers’ underlying involvement that govern both application usage and advertisement response. Interestingly, consumers’ responses to advertisements follow an inverted-U shape where consumers are most likely to respond to advertisements in a medium state of involvement. Consumers’ advertisement responses are also subject to a contextual effect. For example, consumers are more likely to respond to advertisements when they use Entertainment apps compared with other apps. Our simulation indicates that incorporating mobile usage information, such as a temporal state of involvement and contextual effects at the individual level (viewing history), can significantly improve the effectiveness of targeting strategies. For instance, incorporating contextual effect and multitasking can increase performance by as much as 21.2%. This improvement can be further enhanced with the help of the LSTM autoencoder to address the long-range correlations in HMM. We are the first to connect consumers’ mobile application usage with their in-app ad response.
Nov. 20, 2024
Annie Qu :
4 p.m. in 636 SEO
Abstract
Recent advances in dynamic treatment regimes (DTRs) provide powerful optimal treatment searching algorithms, which are tailored to individuals’ specific needs and able to maximize their expected clinical benefits. However, existing algorithms could suffer from insufficient sample size under optimal treatments, especially for chronic diseases involving long stages of decision-making. To address these challenges, we propose a novel individualized learning method which estimates the DTR with a focus on prioritizing alignment between the observed treatment trajectory and the one obtained by the optimal regime across decision stages. By relaxing the restriction that the observed trajectory must be fully aligned with the optimal treatments, our approach substantially improves the sample efficiency and stability of inverse probability weighted based methods. In particular, the proposed learning scheme builds a more general framework which includes the popular outcome weighted learning framework as a special case of ours. Moreover, we introduce the notion of stage importance scores along with an attention mechanism to explicitly account for heterogeneity among decision stages. We establish the theoretical properties of the proposed approach, including the Fisher consistency and finite-sample performance bound. Empirically, we evaluate the proposed method in extensive simulated environments and a real case study for COVID-19 pandemic.
Jan. 15, 2025
Kentaro Takeda :
4 p.m. in Zoom
Abstract
The primary purpose of a dose-finding trial for novel anticancer agents is to identify an optimal dose (OD), defined as the tolerable dose that has adequate efficacy in unpredictable dose-toxicity and dose-efficacy relationships. The FDA project Optimus reforms the paradigm of dose optimization and recommends that dose-finding trials compare multiple doses to generate these additional data at promising dose levels. The backfill is helpful in settings where the efficacy of a drug does not always increase with the dose level. More information is available at these doses by backfilling patients at lower doses while the trial continues to explore higher doses. This paper proposes a Bayesian optimal interval design using efficacy and toxicity outcomes that allows patients to be backfilled at lower doses during a dose-finding trial while prioritizing the dose-escalation cohort to explore a higher dose. A simulation study shows that the proposed design, the BF-BOIN-ET design, has advantages compared to the other designs in terms of the percentage of correct OD selection, reducing the sample size, and shortening the duration of the trial in various realistic settings.
Feb. 12, 2025
Radoslav Harman :
4 p.m. in Zoom
Abstract
The field of optimal experimental design has traditionally focused on “approximate” designs, which specify a finite set of experimental conditions along with the proportions of trials allocated to each condition. The main advantage of approximate designs is that they allow the use of powerful theoretical and numerical tools from convex optimization.
In practical applications, however, “exact” experimental designs are required. These designs determine a finite set of experimental conditions for conducting the trials. Although an exact design can often be derived from an approximate design using rounding algorithms, such procedures typically yield suboptimal results.
In this talk, I will first define and review the integer optimization problem underlying optimal exact design for statistical models with uncorrelated observations, emphasizing its theoretical and computational complexity. Next, I will survey various approaches for computing optimal exact designs numerically. In particular, I will highlight popular exchange methods and other heuristic strategies. I will also discuss methods based on mixed-integer mathematical programming formulations, including the latest developments in the field. Finally, I will illustrate these methods on several challenging problems involving exact optimal designs under non-standard experimental constraints.
Feb. 26, 2025
Huiling Liao :
4 p.m. in 636 SEO
Abstract
Motivated by the challenges in detecting extremely rare failures for sophisticated specifications in circuit design, we consider the problem of detecting regions of interest (ROIs) that consist of specifications with the value of a complex target function for the system performance being below or above a certain pre-specified threshold. Though Bayesian optimization (BO) has been applied to this problem, it is not effective in identifying multiple ROIs as it was originally designed for global optimization and tends to focus on searching the area where the global optimum is most likely to be. In this work, we propose a sampling strategy for fast ROI detection within a limited number of target function evaluations. The sampling distribution is designed so that the probability of a specification being sampled is proportional to the corresponding value of the acquisition function. Such an acquisition-guided sampling algorithm promotes a wider search of the sample space and a simpler incorporation of different criteria to determine the specifications to be evaluated next. To further improve the performance, we propose a new design of the acquisition function and two modifications of existing acquisition functions. Numerical studies on synthetic functions and a real-world circuit design application demonstrate that the proposed method can enjoy a stronger exploration ability provided by sampling and achieve faster ROI detection with higher coverage.
March 19, 2025
Tong Chen :
4 p.m. in Zoom
Abstract
In regression models fitted to data from complex survey designs, sampling weights often incorporate non-essential variation, inflating variance estimates. Stabilised weights mitigate this issue by adjusting sampling weights to account for variation explained by covariates. We evaluate the performance of optimal stabilised weights and propose combining the stabilised weights estimator with generalised raking, a class of efficient design-based estimators. This combination improves efficiency by reducing unnecessary weight variation and leveraging information from auxiliary variables. We show this combination can be implemented using the standard statistical package that handles two-phase samples and generalised raking. Simulation studies demonstrate that the proposed estimator enhances precision under realistic two-phase designs, though efficiency gains may be limited in highly informative designs.
April 2, 2025
Ryan Lekivetz :
4 p.m. in 636 SEO
Abstract
Testing statistical software is an extremely difficult task. What is more, for many statistical packages, the developer and test engineer are one and the same, may not have formal training in software testing techniques, and may have limited time for testing. This makes it imperative that the adopted testing approach is both efficient and effective and, at the same time, it should be based on principles that are readily understood by the developer. As it turns out, the construction of test cases can be thought of as a designed experiment (DOE). This talk provides a treatment of DOE principles applied to testing statistical software and includes other considerations that may be less familiar to those developing and testing statistical packages.
April 9, 2025
Yao Li :
4 p.m. in 636 SEO
Abstract
As machine learning models become increasingly integrated into distributed and language-intensive applications, ensuring their integrity against backdoor attacks is paramount. This talk presents two defense strategies that target vulnerabilities in federated learning and large language models (LLMs). The first part introduces Trusted Aggregation (TAG), a robust defense mechanism for federated learning that leverages a small validation set to estimate permissible updates and filter out malicious contributions. TAG effectively mitigates backdoor risks while preserving task accuracy, even when up to 40% of client updates are adversarial. The second part addresses the threat of syntactic textual backdoor attacks in LLMs. We propose a novel token substitution strategy that alters semantic content while preserving syntactic structures, enabling the detection of both syntax-based and token-based triggers.
April 16, 2025
Nan Xi :
4 p.m. in 636 SEO
Abstract
Combination drug therapies hold significant promise in enhancing treatment efficacy, particularly in fields such as oncology, immunotherapy, and infectious diseases. However, designing clinical trials for these regimens poses unique challenges due to multiple hypothesis testing, shared control groups, and overlapping treatment components that induce complex correlation structures. In this work, we develop a novel statistical framework tailored for early-phase translational combination therapy trials, with a focus on platform trial designs. Our methodology introduces a generalized Dunnett’s procedure that controls false positive rates by accounting for the correlations between treatment arms. Additionally, we propose strategies for power analysis and sample size optimization that leverage preclinical data to estimate effect sizes, synergy parameters, and inter-arm correlations. Simulation studies demonstrate that our approach not only controls various false positive metrics under diverse trial scenarios but also informs optimal allocation ratios to maximize power. A real-data application further illustrates the practical integration of translational preclinical insights into the clinical trial design process. Overall, our framework provides practical and statistically robust guidance for the design of early-phase combination therapy trials, enhancing the efficiency of the bench-to-bedside transition.
April 23, 2025
Gilbert W. Bassett :
4 p.m. in 636 SEO
Abstract
The VIX--volatility index--is a number. It is transmitted to the world every 15 seconds from downtown Chicago, viewed right outside the window from UIC, at Cboe. It features Dispersion, Probability, Fear, and other topics of interest to statistics students. It is not related to the Black Scholes model, its volatility is not Variance, and the related Volatility of Volatility (VVIX) measures the Fear of Fear. An introduction to the VIX accessible to statistics students is presented via the Science Fictional Options Universe (SFOU) wherein, among other things, Probability does not exist: Never Happened. Not subjective, physical, or frequentist; no dice. People understand that the future is uncertain and have a sense about what is more or less likely, but "likely" in the SFOU is an informal concept, you know what I mean. In our universe it is like Hygge.
*Disclosure: I am on the board of directors at the Cboe Futures Exchange that produces the VIX.
Sept. 10, 2025
TBA :
4:15 p.m. in 636 SEO
Sept. 24, 2025
Hongyuan Cao :
4:15 p.m. in 636 SEO
Abstract
Testing composite null hypotheses is fundamental to many scientific applications, including mediation and replicability analyses, and becomes particularly challenging in high-throughput settings involving tens of thousands of features. Existing high-dimensional composite null hypotheses testing often ignores the dependence structure among features, leading to overly conservative or liberal results. To address this limitation, we develop a four-state hidden Markov model (HMM) for bivariate $p$-value sequences arising from two-study replicability analysis. This model captures local dependence among features and accommodates study-specific heterogeneity. Based on the HMM, we propose a multiple testing procedure that asymptotically controls the false discovery rate (FDR). Extending this framework to more than two studies is computationally intensive, with complexity growing exponentially in the number of studies $n.$ To address this scalability issue, we introduce a novel e-value framework that reduces computational complexity to quadratic in $n,$ while preserving asymptotic FDR control. Extensive simulations demonstrate that our method achieves higher power than existing approaches at comparable FDR levels. When applied to genome-wide association studies (GWAS), the proposed approach identifies novel biological findings that are missed by current methods.
Oct. 1, 2025
Qiong Zhang :
4:15 p.m. in Zoom
Abstract
A/B testing is an effective method to assess the potential impact of two treatments. For A/B tests conducted by IT companies like Meta and LinkedIn, the test users can be connected and form a social network. Users’ responses may be influenced by their network connections, and the quality of the treatment estimator of an A/B test depends on how the two treatments are allocated across different users in the network. In this talk, I will discuss optimal design criteria based on some commonly used outcome models, under assumptions of network-correlated outcomes or network interference. I will show that the optimal design criteria under these network assumptions depend on several key statistics of the random design vector. I will discuss a framework to develop algorithms that generate rerandomization designs meeting the required conditions of those statistics. I further talk about asymptotic distributions to guide the specification of algorithmic parameters and validate the proposed approach using both synthetic and real-world networks.
Oct. 22, 2025
Dr. Hongwen Guo :
4:15 p.m. in Zoom
Abstract
In the era of digital assessments, large-scale educational data—such as that from NAEP—offers unprecedented opportunities for insight into student learning skills. Yet, the complexity and volume of this data often outpace traditional statistical approaches, calling for a fusion of statistical rigor, data science innovation, and AI-driven modeling.
This talk explores a research initiative at ETS, supported by the Gates Foundation, that helps to transform multi-source NAEP data (response, process, and behavioral) into actionable insights for educators. We will discuss how statistics and data science form the foundation for extracting meaningful patterns, visualizing complex data, and how human-centered AI enables scalable, interpretable feedback with subject-matter experts and teachers.
The presentation will also reflect on the speaker’s own professional evolution—from classical statistics to data science and to AI applications in the education measurement field —highlighting the synergies between these disciplines.
Faculty and graduate students interested in statistical modeling, educational measurement, and AI applications are invited to join the discussion.
Oct. 29, 2025
Yan Sun :
4:15 p.m. in 636 SEO
Abstract
Precision medicine is the future of drug development, and subgroup identification plays a critical role in achieving the goal. In this presentation, we propose a powerful end-to-end solution squant (available on CRAN) that explores a sequence of quantitative objectives. The method converts the original study to an artificial 1:1 randomized trial, and features a flexible objective function, a stable signature with good interpretability, and an embedded false discovery rate (FDR) control. We demonstrate its performance through simulation and provide a real data example.
Nov. 5, 2025
Heejong Bong :
4:15 p.m. in 636 SEO
Abstract
In network settings, interference between units makes causal inference more challenging as outcomes may depend on the treatments received by others in the network. Typical estimands in network settings focus on treatment effects aggregated across individuals in the population. We propose a framework for estimating node-wise counterfactual means, allowing for more granular insights into the impact of network structure on treatment effect heterogeneity. We develop a doubly robust and non-parametric estimation procedure, KECENI (Kernel Estimator of Causal Effect under Network Interference), which offers consistency and asymptotic normality under network dependence. The utility of this method is demonstrated through an application to microfinance data, revealing the node-wise impact of network characteristics on treatment effects.
Nov. 19, 2025
Dogyoon Song :
4:15 p.m. in Zoom
Abstract
Regression adjustment is a classical technique in causal inference that leverages covariates to improve precision of estimators in randomized controlled trials (RCTs) and to adjust for confounding in observational studies. While well-understood in low-dimensional settings, its behavior in modern high-dimensional regimes---where the number of covariates may be comparable to or even exceed the number of observations---remains underexplored. In particular, existing theoretical results are largely asymptotic, often rely on residual-based arguments, and provide limited insights into finite-sample inference especially when $p>n$.
In this talk, we revisit regression adjustment for the average treatment effecting (ATE) estimation under complete randomization with many covariates, in a design-based, finite-population framework, via two vignettes. First, we introduce a novel theoretical perspective on the asymptotic properties of regression adjustment through a Neumann-series decomposition, yielding a refined analysis in the $p<n$ regime. Specifically, for ordinary least squares (OLS) regression adjustment, we show that the degree-$d$ Neumann-corrected estimator is asymptotically normal when $p^{d+3}(\log p)^{d+1}=o(n^{d+2})$. This result strictly enlarges the previously reported admissible growth of $p = o(n^{1/2})$ or $p = o(n^{2/3})$ with a single de-biasing step. Second, we present a non-asymptotic analysis of the regression-adjusted ATE estimators that is valid in both $p<n$ and $p>n$ settings. Leveraging concentration of measure tools, we quantify uncertainty without relying on classical asymptotic variance estimation, and further control the design bias of estimators via Stein's method of exchangeable pairs. Time permitting, we will discuss potential extensions and ongoing work.
Dec. 3, 2025
Sixia Chen :
4:15 p.m. in Zoom
Abstract
Non-probability samples are prevalent in various fields, such as biomedical studies, educational
research, and business investigations, owing to the escalating challenges associated with
declining response rates and the cost-effectiveness and convenience of utilizing such samples. However, relying on naive estimates derived from non-probability samples, without adequate adjustments, may introduce bias into study outcomes. Addressing this concern, data integration methodologies, which amalgamate information from both probability and non-probability samples, have demonstrated effectiveness in mitigating selection bias. Nonetheless, the efficacy of these methods hinges upon the assumptions underlying the models. This paper introduces innovative and robust data integration approaches, notably a semi-parametric quantile
regression-based mass imputation approach and a doubly robust approach that integrates a non-
parametric estimator of the participation probability for non-probability samples. Our proposed
methodologies exhibit greater robustness compared to existing parametric approaches,
particularly concerning model misspecification and outliers. We consider both missing at random
and not missing at random scenarios. Theoretical results are established, including variance estimators for our proposed estimators. Through comprehensive simulation studies and real-
world applications, our findings demonstrate the promising performance of the proposed
estimators in facilitating valid statistical inference. This research contributes to the advancement
of robust methodologies for handling non-probability samples, thereby enhancing the reliability
and validity of research outcomes across diverse domains.
Feb. 25, 2026
Jie Jian :
4:15 p.m. in 636 SEO
Abstract
Detecting dependence structures in international trade—such as persistent exporter–importer affinities, and supply-chain clustering—often relies on latent variable models that summarize high-dimensional trading flows. We propose a novel Bayesian non-negative tensor factorization for large, sparse, nonnegative trading tensors with excess zeros and continuous positive measurements. We target settings with millions of entries and extreme sparsity. Each entry follows a spike-and-slab model: a point mass at zero coupled with a gamma–Poisson construction that yields a low-rank nonnegative decomposition via gamma latent factors. The framework provides interpretable mode-specific components and principled uncertainty quantification.
March 11, 2026
Lingjie Ma :
4:15 p.m. in 636 SEO
Abstract
It is well known that asset returns usually do not follow a normal distribution, rather, they have long and fat tails. This paper focuses on the quantile portfolio methodology, which considers the whole distribution of asset returns and employs expected loss as a risk measurement. In particular, we explore statistical properties of tau risk and propose related theories of quantile portfolio optimization. We also introduce portfolio performance terms for the quantile portfolio framework.
March 18, 2026
Dr. Xiao Zhang :
4:15 p.m. in Zoom
Abstract
Probit models have been prominent tools to analyze binary/ordinal data, but the computational complexity of maximum likelihood functions presents challenges in their usage. Furthermore, the model identification necessitates the covariance matrix of the latent multivariate normal variables to be a correlation matrix, which brings a rigorous task to develop efficient Markov chain Monte Carlo (MCMC) sampling methods. Data augmentation has been inevitable explored for both identifiable univariate and multivariate probit models. Particularly, it is well-known that parameter-expanded data augmentation (PX-DA) based on non-identifiable models accelerates the convergence and improves the mixing of MCMC components. However, comprehensive investigation has seldom been undertaken, and various algorithms due to incorrectly constructed non-identifiable models further bring obstacles to develop efficient MCMC sampling methods. We tackle this issue by constructing correct non-identifiable models and develop PX-DA algorithms to estimate both univariate and multivariate probit models. Our investigation exhibits that the proposed PX-DA algorithms advance the performance of MCMC sampling considerably and illustrates the essentials of using PX-DA, especially for data with large sample sizes.
April 1, 2026
Hsin-Hsiung Huang :
4:15 p.m. in 636 SEO
Abstract
We develop a new Fréchet sufficient dimension reduction (FdSDR) framework tailored for high-dimensional functional data with complex, metric-space-valued responses. Our main contribution is a low-rank distance covariance criterion that enables scalable, model-free identification of low-dimensional predictor structures while capturing nonlinear dependence. The proposed method is computationally efficient in high dimensions and avoids restrictive distributional assumptions. We establish theoretical guarantees and demonstrate its effectiveness through simulations and real data, providing a practical and flexible approach for modern functional data analysis.