Skip to main content

Statistics and Data Science Seminar : Past Events

Past Seminars

The following seminars have already happened, you may instead view upcoming seminars in this series.

Aug. 31, 2005

Professor Sujit Kumar Mitra: An Appreciation

Professor Dibyen Majumdar : 3 p.m. in SEO 512
Abstract Professor Sujit Kumar Mitra was a luminary in the world of statistics and linear algebra. The clarity and brilliance of his results and proofs have been seldom matched. I will attempt an appreciative reflection on his work through examples on asymptotics of the contingency chi-square, characterization of the Wishart distribution, MINQUE theory, matrices and the g-inverse. The talk will be at the graduate student level.

Sept. 14, 2005

High frequency financial data and the hidden semimartingale model

Professor Lan Zhang : 3 p.m. in SEO 512
Abstract The availability of high frequency data for financial instruments has opened the possibility of accurately determining volatility in small time periods, such as one day. Recent work on such estimation indicates that it is necessary to treat the data with a hidden semimartingale model, typically by the addition of measurement error. We review the emerging theory on this subject, including two- and multiscale sampling.

Sept. 21, 2005

Locally D-optimal Designs Based on Combined Emax Model and PK Compartmental Models

Xin Fang : 3 p.m. in SEO 512
Abstract Both pharmacodynamic (PD) Emax models and pharmacokinetic (PK) one-compartment models have been well studied and widely used in pharmaceutical industry to obtain the ADME (absorption, distribution, metabolism and excretion) and evaluate the efficiency of a drug candidate. However, the implementation of experimental designs on Emax models is difficult since concentrations of the drug in blood are uncontrollable. In this talk, two models that combine the PD and PK models are discussed to address this problem. With the current models, one can obtain better estimations of the parameters of the ADME and efficiency of a drug through implementing a locally D-optimal design. A class of robust designs is also investigated for comparison. Simulation results show that the robust designs have better statistical properties when sample size is large such as in clinical trial III, while the locally D-optimal designs are more proper when sample size is small as in clinical trials I & II.

Sept. 28, 2005

Non-Inferiority Testing in Thorough QT/QTc Studies

Professor Balakrishna Hosmane : 3 p.m. in SEO 512
Abstract Evaluation of new drugs for unwanted effects on electrical properties of the heart is receiving heightened attention from pharmaceutical companies and regulatory agencies. This attention arises from recent scientific research that links drug effects in cellular ion channels to changes in electrical characteristics of the electrocardiogram (ECG) that predict clinically important cardiac arrhythmias. An assessment of non-inferiority of the higher dose to placebo is performed in four-period cross-over study as well as four-group parallel study by the union- intersection test within the framework of a linear mixed effects analysis. For the purposes of planning such a study, the joint distribution of the estimate of the difference in means of high dose of investigational drug and the placebo was derived. The power of thorough QT/QTc study evaluated using the joint distribution and the simulation study were quite close.

Oct. 5, 2005

A Two-phase Probability Sampling Process

Payal Pagni : 3:30 p.m. in SEO 512
Abstract We will discuss a two-phase sampling situation where the first phase consists of demographic collection and the second phase consists of a controlled selection of the sample.

Oct. 12, 2005

Finding Optimal Policies in Markov Decision Processes and Stochastic Games

Jaime Brugueras : 3 p.m. in SEO 512
Abstract Markov decision processes (MDP) are stochastic processes that describe the evolution of dynamic systems controlled by sequences of decisions or actions. Different paths of the system lead to associated economic consequences; the ultimate aim is to take those actions that optimize a certain criterion. I will review the mathematical model of such processes, give real-life examples, and describe the well-known algorithms for finding optimal policies. Stochastic games are a natural generalization of MDP to the case of two or more controllers. Existence of finite algorithms for finding optimal stationary policies is in general an open problem. I consider a special class of stochastic games, those with perfect information, which can be solved via a finite algorithm.

Oct. 19, 2005

Design of Accelerated Thermal Degradation Experiments

Dr. William R. Porter : 3:30 p.m. in SEO 512
Abstract Accelerated testing of product quality attributes is an important tool used in the development of new products. Many products fail to meet quality specifications after storage for prolonged times due to degradation of the materials used to manufacture the product. The time to failure is the shelf life of the product. Thermal stress is often used as an accelerant to promote degradation in a controlled manner to predict shelf life. The results of stress degradation experiments are evaluated using the nonlinear Arrhenius kinetic model. Historical methods of fitting the Arrhenius model to thermal degradation experimental data, particularly the Garrett approach, will be discussed. Optimal design of thermal degradation experiments from a theoretical basis has been described in the literature but many practical pitfalls remain. Nonlinear model design requires knowledge of the model parameters, crude estimates of which can be obtained using a pilot experiment. Monte Carlo evaluation of the design of a pilot experiment will be discussed, along with opportunities for future development of practical designs for refining parameter estimates.

Oct. 26, 2005

An Unified Approach for Measuring Agreement

Wenting, Wu : 3:30 p.m. in SEO 512
Abstract Assessing agreement plays an important role in assessing the acceptability of a new or generic process, methodology, and formulation in areas of laboratory performance, instrument or assay validation, method comparisons, and individual bioequivalence. An introduction of current existing methods for measuring agreement will be given. After that, an unified approach for measuring agreement of k readers each with multiple readings will be introduced. When each reader has only one reading, this mehtod degenerates into the conventional overall agreement index. When there are only two readers, it degenerates into the conventional concordance correlation coefficient (CCC). When data are ordinal, it degenerates into the weighted kappa coefficient. When data are binary, it degenerates into the kappa coefficient. The approach uses generalized estimation equations (GEE) to model various functions of variance components. Some simulation results will be presented and followed by an example.

Nov. 2, 2005

Generalized Linear Models For the Covariance Matrix of Longitudinal Data

Mohsen Pourahmadi : 3:30 p.m. in SEO 512
Abstract We survey the progress made in modelling covariance matrices from the perspective of generalized linear models (GLM) and show how one can move beyond the use of the identity and logarithmic link functions, and prespecified structures. Observing that most time-domain models (ARMA, state-space,....) in time series analysis are means to diagonalize a Toeplitz covariance matrix via a unit lower triangular matrix (Cholesky decomposition), we discuss the distinguished role of the Cholesky decomposition in providing a systematic and data-based procedure for formulating and fitting parsimonious models for general covariance matrices guaranteeing the positive-definiteness of the estimates. Pulling together some techniques from regression and time series analyses provide the necessary tools for the procedure which reduces the unintuitive task of modelling covariance matrices to that of a sequence of regression models. The procedure is illustrated using a real longitudinal dataset.Once a bona fide GLM framework for modelling covariances is found, its bayesian, nonparametric, generalized additive and other extensions can be developed in direct analogy with the respective extensions of the traditional GLM.

Nov. 9, 2005

Critical Sets in Mutually Orthogonal Latin Squares

Rita Saha Ray : 3:30 p.m. in SEO 512
Abstract A critical set consists of the minimum information needed to recreate a combinatorial structure uniquely. To date, very few results on critical sets for a set of Mutually Orthogonal Latin Squares [MOLS] are known. In the present talk, we consider k Mutually Orthogonal Cyclic Latin Squares of order n, n odd, and obtain bounds on the possible sizes of the minimal critical sets. For n = 7, we consider a complete set of MOLS and exhibit a minimal critical set, improving upon the bound reported in Keedwell (1997). The problem is also addressed for a pair of MOLS of odd order n, n ≥ 9. Critical sets achieving the proposed bound are obtained for n = 9 and 15. (This is a joint work with Avishek Adhikari and Jennifer Seberry)

Nov. 16, 2005

Bayesian applications in non clinical pharmaceutical statistics

Dr. David LeBlond : 3:30 p.m. in SEO 512
Abstract This talk includes about 3 topics. 1. Predicting manufacturing failure rates (one way random modeling); 2. Estimation of shelf life from accelerated stability studies, a followup to Bill Porter's talk (non linear modeling); 3. Comparison of 2 analytical methods, both subject to error (Errors in variables/ latent variables modeling). Also, a little of WinBUGS demo will be given.

Nov. 23, 2005

Maximum Likelihood Estimation of Point Scatterers for Computational Time-reversal Imaging

Gang Shi : 3:30 p.m. in SEO 512
Abstract We present a statistical framework for the fixed-frequency computational time-reversal imaging problem assuming point scatterers in a known background medium. Our statistical measurement models are based on the physical models of the multistatic response matrix, the distorted wave Born approximation and Foldy-Lax multiple scattering models. We develop maximum likelihood (ML) estimators of the locations and reflection parameters of the scatterers. Using a simplified single-scatterer model, we also propose a likelihood time-reversal imaging technique which is suboptimal but computationally efficient and can be used to initialize the ML estimation. We generalize the fixed-frequency likelihood imaging to multiple frequencies, and demonstrate its effectiveness in resolving the grating lobes of a sparse array. This enables to achieve high resolution by deploying a large-aperture array consisting of a small number of antennas while avoiding spatial ambiguity. Numerical and experimental examples are used to illustrate the applicability of our results.

Nov. 30, 2005

Estimation of a Normal Mean with Known CV

Weiya Zhang : 3:30 p.m. in SEO 512
Abstract Historically, this problem has been encountered in statistical quality control situations and has since been extensively studied in statistical literature on theory & applications. Notable contributors in this fascinating topic are Wolfowitz (AMS, 1957), Searls (1963, JASA1964), Hendricks (JASA, 1964), Khan (JASA, 1968), Azen & Reed (Technoterics, 1973), Gleser & Heely (JASA, 1976), Sen (BJ, 1979), Soofi & Gokhale (CSDA, 1991), Guo & Pal (CSAB, 2003), Chaturvedi & Tomer (Statistics, 2003) and Singh & Mathur (JSPI, 2005). Recently we came across a Review Article by Anis (2005). I will present the review article and supplement the study with some of our own research findings.

Jan. 18, 2006

Logistic Regression Trees

Professor Kin-Yee Chan : 3:30 p.m. in SEO 512
Abstract Logistic regression is a powerful technique for fitting models to data with a binary response variable, but the models are difficult to interpret if collinearity, nonlinearity, or interactions are present. Besides, it is hard to judge model adequacy since there are few diagnostics for choosing variable transformations and no true goodness-of-fit test. To overcome these problems, we propose to fit a piecewise (simple,multiple or stepwise) linear logistic regression model by recursively partitioning the data and fitting a different logistic regression in each partition. This allows nonlinear features of the data to be modeled without requiring variable transformations. Trend-adjusted chi-square tests are used to control bias in variable selection at the intermediate nodes. This protects the integrity of inferences drawn from the tree structure. The binary tree that results from the partitioning process is pruned to minimize a cross-validation estimate of the predicted deviance. This obviates the need for a formal goodness-of-fit test. Our algorithm, called "LOTUS", is compared with standard stepwise logistic regression and two well-known classification tree algorithms (QUEST and C4.5) on 13 real datasets, with several containing tens to hundreds of thousands of observations. Results will be presented at this talk.

Feb. 6, 2006

Statistical Learning Problems in Molecular Biology

Professor Probal Chaudhuri : 4 p.m. in SEO 636
Abstract Statistical learning problems arise in the study of molecular evolution as well as in the development of gene predictors. Problems can be unsupervised, supervised or partially supervised in nature depending on the situation. This talk will discuss the use of oligonucleotide distributions in DNA sequences in solving such problems. Some related probabilistic models for DNA sequences will also be discussed. In particular, a stochastic replication model for biological sequences will be introduced that generalizes standard hidden Markov models.

March 1, 2006

Statistics in Clinical Drug Development

Dr. Jeen Liu : 3:30 p.m. in SEO 512
Abstract This presentation will be an overview of the role of statistics and statisticians in the drug development process. I will present an overview of the pharmaceutical industry, drug development process, and the clinical research organizations. That will be followed by the challenges for statisticians in the industry. The current statistical issues of interest and some examples will also be provided.

March 8, 2006

Impact of Missing Data on Building Prognostic Models and Summarizing Models Across Studies

Dr. Mahtab Munshi : 3:30 p.m. in SEO 512
Abstract We examine the impact of missing data in two settings, the development of prognostic models and the addition of new risk factors to existing risk functions. Most statistical software presently available performs complete case analysis, wherein only participants with known values for all of the characteristics being analyzed are included in model development. Missing data also impacts the summarization of evidence amongst multiple studies using meta-analytic techniques. As we progress in medical research, new covariates become available for studying various outcomes. While we want to investigate the influence of new factors on the outcome, we also do not want to discard the historical datasets that do not have information about these markers. We investigate different methods to estimate parameters for a model when some of the covariates are missing. These methods include likelihood-based inference for the study-level coefficients and likelihood based inference for the logistic model on the person-level data. We compare the results from our methods to the corresponding results from complete case analysis. We focus our empirical investigation on a historical example, the addition of high-density lipoproteins to existing equations for predicting death due to coronary heart disease. We verify our methods through simulation studies on this example.

March 15, 2006

Some Aspects of Statistical Inference on a Normal Mean with Known Coefficient

Weiya Zhang, Ph.D. Candidate : 2:30 p.m. in SEO 512
Abstract Inference on a normal mean with known CV is intricate since the distribution does not admit of a complete sufficient statistic. Consequently, no umvue exists for the mean. As a result, there have been many attempts to suggest biased estimators which are functions of the sample mean and the sample sd. There is an extensive literature on this fascinating topic. However, the mean parameter is assumed to be positive-valued. If we allow the parameter space to include negative values of the mean as well, the problem of unbiased estimation of the mean in terms of the sample standard deviation becomes intriguing. Starting with the framework of n(>1) i.i.d. observations from a normal population with known CV, we offer (i) an analytical expression for exact unbiased estimator of the mean in terms of the sample sd and the sign function of the sample mean; (ii) an analytical expression for the best linear combination of the estimator in (i) and the sample mean as an unbiased estimator for the mean, along with exact expression for the variance of this linear combination; (iii) a study of asymptotic normality of the linear combination in (ii) as well as its behavior in small samples; (iv) confidence interval for the mean based on a variation of the best linear combination in (ii); (v) comparison of traditional confidence interval for the mean and the one suggested in (iv); (vi) improved fixed width confidence interval for the mean.

March 29, 2006

Statistician's Role in a Pharmaceutical Company and Common Issues in Clinical Trial Design

Dr. Weining Z Robieson : 3:30 p.m. in SEO 512
Abstract Clinical trial design issues from statistical point of view will be presented. Statisticians' role in a pharmaceutical company will be illustrated by description of statisticians' job responsibilities. The following issues will be discussed: 1)Handling of missing data; 2)Adjustment for baseline covariates; 3)Interim analysis; 4)Multiple comparison; 5)Adaptive design.

Optimal crossover designs

Professor Min Yang : 10 a.m. in SEO 512
Abstract Crossover designs, where experimental subjects are used in two or more (p) periods for the purpose of evaluating and studying two or more (t) treatments, originated from agricultural studies and have proven widely effective in a variety of fields, especially in phase I and phase II pharmaceutical clinical trials. The rigorous study of these designs and their optimality and efficiency has a history of more than 3 decades. In this talk, we will review and study the optimality, efficiency, and robustness of crossover designs under the following two different situations (i) all treatment comparisons are equally important and (ii) for comparing several test treatments to a control treatment. Two algorithms, both guided by these efficiencies and results from optimal design theory, are proposed for obtaining efficient designs under the various models.

April 5, 2006

Investigating the Categories for Cholesterol and Blood Pressure for Risk Assessment of Death due to Coronary Heart Disease

Dr. Billy Franks : 3:30 p.m. in SEO 512
Abstract Many characteristics for predicting death due to coronary heart disease are measured on a continuous scale. These characteristics, however, are often categorized for clinical use. We suggest a systematic approach to determine the best categorizations of systolic blood pressure and cholesterol level for use in identifying individuals who are at high risk for death due to coronary heart disease. We also compare these data derived categories to those in common usage. A version of Classification And Regression Trees (CART) that can be applied to censored survival data will be used to identify categories in multiple data sets. The collection of categories will then be used to identify major cut-points, which are common in all of the data sets by using kernel density estimation.

April 12, 2006

Hypothesis Testing, Power and Sample Size Determination for Between Group Comparisons in fMRI Experiments

Professor Dulal K. Bhaumik : 3:30 p.m. in SEO 512
Abstract Modern methods for imaging the human brain, such as functional magnetic resonance imaging (fMRI)present a range of challenging statistical problems. In this talk, we will look at a number of possible models for the analysis of fMRI data from multiple subjects, and develop tests that lead to calculations of power and sample size for between group comparisons. Sample size calculations are particularly critical for neuroscientists who use these new techniques, since each subject is expensive to image.

April 14, 2006

Exponentially Weighted Moving Average Methods for Detection of Change Point for Event Rates

Yuping Dong : 3:30 p.m. in SEO 512
Abstract The exponentially weighted moving average methods (EWMA) are one of the statistical surveillance methods commonly studied in statistical process control literature. The EWMA methods are mainly used to monitor the mean of the distribution of a continuous quality measure. In this talk, we present a way to extend the EWMA procedure to the case of a positive shift in the incidence rate per exposure unit of a Poisson process. Three types of EWMA methods, EWMAe, EWMAa1 and EWMAa2, are constructed, all with an alarm statistic, which is an exponentially weighted moving average of the observations per exposure unit. Analytical bounds for different measures of evaluation, suitable in different types of applications, are provided such as the expected delay, the average run length to an alarm and the probability of successful detection, to give a broad picture of the features of the methods. Results from a simulation study are presented both for a fixed average run length to the first false alarm and a fixed probability of a false alarm.

April 26, 2006

Quantile Regression Modeling and Estimation for Paired Comparisons

Gib Bassett : 3:30 p.m. in SEO 512
Abstract After a brief introduction to quantile regression and modeling the talk will consider paired comparisons. The context is rating and ranking sports teams based on game outcomes. The quantile regression approach includes the standard model as a special case, while allowing for a richer set of possible relationships between teams and outcomes. Compared to models that focus on one part of a distribution, the quantile approach expresses relationships that depend on the different parts of a distribution. We can have "A better than B" based on the expected outcome, while at the same time B is more likely to win the game. The ratings are defined as handicaps that make handicap-adjusted outcomes of games evenly matched (where "equally matched" depends on which property of the outcome distribution is to be equalized). We consider connections to point spreads and odds wagering as well as the roundness of a round robin.

Sept. 20, 2006

Some remarks on the use of trees in statistical models

Peter McCullagh : 3:30 p.m. in SEO 636
Abstract Trees arise naturally in the study of the evolution of a population. Branching processes are the natural tool in the forward direction, and coalescent processes are natural for the study of ancestral relationships or lineages in reverse time. A tree has a natural graphical representation, but for some purposes a matrix representation is also useful. In statistical work, a similarity matrix is a covariance matrix generated by additive common factors with independent components. The set of similarity matrices also coincides with the set of fragmentation trees. Some issues arising in the use of structured covariance matrices of this sort will be discussed.

Sept. 27, 2006

ASA Membership Social

S. Hedayat : 3:30 p.m. in SEO 512

Oct. 4, 2006

Loglinear residual tests of Moran' I autocorrelation and their applications to Kentucky breast cancer data

Tonglin Zhang : 3:30 p.m. in SEO 512
Abstract Moran's I is the most widely used and the most frequently cited test statistic in spatial statistical literature. This research bridges the permutation test of Moran's I to the residuals of a loglinear model under the asymptotic normality assumption. It provides the versions of Moran's I based on Pearson residuals ( ) and deviance residuals ( ) so that they can be used to test for spatial clustering while at the same time account for potential covariates and heterogeneous population sizes. Our simulations showed that both and are effective to account for heterogeneous population sizes. The tests based on and are applied to a set of loglieanr models for early stage and late-stage breast cancer with socioeconomic and access-to-care data in Kentucky. The results showed that socioeconomic and access-to-care variables can sufficiently explain spatial clustering of early stage breast carcinomas, but these factors cannot explain that for the late-stage. For this reason, we used local spatial association terms and located four late-stage breast cancer clusters that could not be explained. The results also confirmed our expectation that a high screening level would be associated with a high incidence rate of early stage disease, which in turn would reduce late-stage incidence rates.

Oct. 18, 2006

FROM DISTRIBUTION-FREENESS TO SEMIPARAMETRIC EFFICIENCY Sixty years of rank-based inference

Professor Marc Hallin : 3 p.m. in SEO 636
Abstract The modern history of ranks in statistics started in 1945 with Frank Wilcoxon's far-reaching four page paper on rank tests for location. Emphasis in 1945 was on distribution-freeness and ease of applications. Since then, under the impulse of such names as Chernoff, Savage, Hodges, Lehmann, Hajek,and Le Cam, rank-based methods have followed the development of contemporary statistics, and turned into a complete body of modern, flexible and powerful techniques. In this talk, we show how this evolution, from distribution- freeness to group invariance and tangent space projections, eventually may reconcile the enemy brothers of statistics---efficiency and robustness.

Oct. 25, 2006

Correction for interarrival time distribution in the Estimation of Poisson Intensity

Dr. Grace L. Yang : 3:30 p.m. in SEO 512
Abstract Occurrence of dead time in recording instruments poses challenging problems in data acquisition, construction of stochastic models and statistical analysis. Well-known examples include the construction of probability models for a paralyzable counter (electron multiplier) and a nonparalyzable counter (e.g., Geiger counter). In this presentation, statistical analysis of recordings from Phase Doppler Interferometry (PDI)is considered. PDI is a non-intrusive technique used to obtain information about spray characteristics in many areas of science, such as liquid fuel spray in combustion, spray coatings, fire suppression and pesticide dispensing. PDI can record the velocity of individual droplets in a spray. However, it will miss some of the droplets because of a recurring presence of dead time. The incompleteness of PDI recordings results in a multimodal interarrival time distribution of droplets. Modeling a spray process as a homogeneous Poisson process, we estimate the spray diffusion rate (Poisson intensity) with correction for dead time under various conditions. The asympotic distribution of the estimates is derived from a strict stationary process. Simulation produced a good agreement between our estimators (in the presence of dead time) and the MLE obtained without dead time. Experimental data from NIST are used for illustration.

Nov. 15, 2006

Bayesian competing risks analysis of cancer survival data from the SEER program

Professor Sanjib Basu : 3:30 p.m. in SEO 512
Abstract The rates of cancers, including age-adjusted mortality and incidence rates, depict a general increase over the last 30 years. These led some to question the success of the war on cancer. The rates of many other competing diseases, on the other hand, have declined. It has been hypothesized that this decline is somewhat responsible for the rise in cancer rates. We consider competing risks analysis of cancer survival data that considers the simultaneous risks of cancer as well as other causes. The cure rate survival models for cancer postulates a fraction of the patients to be cured from cancer. We propose a model that incorporates competing risks and, at the same time, allows a fraction of patients to be cured. We describe Bayesian analysis of this model, discuss both conceptual and methodological issues related to model building and model selection, and consider application in survival data for breast and prostate cancer patients in the SEER registries of the National Cancer Institute (NCI).

Nov. 29, 2006

Stochastic Curtailment under Linear Models

Li Wei : 3:30 p.m. in SEO 512
Abstract Stochastic curtailment, one of the major statistical tools adopted in interim analysis, has attracted more attentions than its competitors such as group sequential procedures for it integrates current data and potential future outcomes in addition to its simplicity in design and implementation. Under this approach, the conditional power, which is the probability of rejecting the null hypothesis at the planned end of the study given the accumulating data, is calculated and the stopping decision is made according to the comparison of this power with a pre-specified threshold. Many procedures with this perspective have been developed for interim analysis. However, possibly for the purpose of statistical convenience, only trials with one or two arms are investigated. Here, we derived an analytic formula for the conditional power under the frame of linear models so that it can be applied to most actual clinical trials in which multiple treatment effects, block effects and covariate effects are all allowed to be considered. The properties of this conditional power is investigated and further our research shows that, unlike the standard power of a regular test for a treatment contrast which depends on unknown parameters only through the contrast itself, the conditional power here fails to have this characteristic in general. A necessary and sufficient condition for the conditional power to depend soly on the interested contrast is provided and some instances are illustrated. Similar arguments can be made about the sufficient statistics for the conditional power. Finally, the results obtained here is applied to an interim analysis performed in a multi-center, randomized, double-blinded, placebo-controlled, parallel group phase II study where centers act as blocks and baseline scores are treated as covariates, resulting in an early termination of the trial and hence a substantial saving in cost.

Jan. 31, 2007

Statistical Auditing in Medicare and Medicaid Fraud Cases

Prof. Klaus Miescke : 4 p.m. in SEO 712
Abstract To estimate the total overpayment on a large number of paid claims, government agencies utilize extrapolation methods that are based on the audit results from a random sample of these claims. Although there is a variety of approaches that are reasonable and statistically valid, some play out better than others in the legal process. For example, the estimator of the total loss must be unbiased by fairness reasons. Using a biased estimator with a smaller MSE would be unacceptable. In this talk we describe the entire audit process, including possible objections from the other side and suggestions on how to respond to them. The theoretical part of the talk is based on W.G. Cochran (1977), Sampling Techniques, 3rd ed., Wiley, NY, and the rest on practical experience. It should be pointed out that it is the loss, not the fraud itself, that can be detected and estimated statistically.

Feb. 14, 2007

Sliced Inverse Moment Regression Using Weighted Chi-Squared

Jie Yang : 3:30 p.m. in SEO 712
Abstract We propose a new class of dimension reduction methods using the first two inverse moments, called Sliced Inverse Moment Regression (SIMR). We develop corresponding weighted chi-squared tests for the dimension of the regression. Basically, SIMR are linear combinations of Sliced Inverse Regression (SIR) and a new method using candidate matrix M_{zz'|y}, which is designed to recover the entire inverse second moment subspace. Theoretically, SIMR, as well as Sliced Average Variance Estimate (SAVE), are more capable of recovering the complete central dimension reduction subspace than SIR and Principle Hessian Directions (pHd). Therefore it can substitute for SIR, pHd, SAVE or any linear combination of them at a theoretical level. Simulation study shows that SIMR using the weighted chi-squared test may have consistently greater power than SIR, pHd, and SAVE.

March 5, 2007

Nonparametric Tests Against Multiple Change Points

Prof. Emad Aly : 3:30 p.m. in SEO 712
Abstract We consider the problem of testing the null hypothesis of no change against the alternative of multiple change points in a series of independent observations. We consider the three cases of testing against the general multiple change point alternative, the ordered multiple change point alternative and the epidemic two change point alternative. We report the asymptotic null distribution of the considered tests. We also give approximations for their limiting critical values.

March 7, 2007

Spline Single-Index Prediction Model

Prof. Lijian Yang : 3:30 p.m. in SEO 712
Abstract For the past two decades, single-index model, a special case of projection pursuit regression, has proven to be an efficient way of coping with the high dimensional problem in nonparametric regression. Applications of single-index model lie in a variety of fields, such as discrete choice analysis in econometrics and dose-response models in biometrics, where high-dimensional regression models are often employed. We investigate the single-index prediction based on weakly dependent sample. The single-index is identified by the best approximation to the multivariate prediction function of the response variable, regardless of whether the prediction function is a genuine single-index function. A polynomial spline estimator is proposed for the single-index coefficients, and is shown to be strongly consistent and asymptotically normal. An iterative program based on Newton-Raphson algorithm is developed. The algorithm is sufficiently fast for the user to analyze large data of high dimension within seconds. Simulation experiments have provided strong evidence that corroborates with the asymptotic theory. Finally, we illustrate our estimation procedure by a gas furnace example.

March 9, 2007

A Small Sample of Data Analysis Problems in Drug Research

Thomas E. Bradstreet, Ph.D. : 2 p.m. in SEO 712
Abstract This example driven presentation outlines some of the basic types of studies traditionally conducted in the pharmaceutical industry. Both animal and human research activities are presented. Special attention is given to statistical analysis issues. Some brief study descriptions and their corresponding data sets can be found at http://www.math.iup.edu/~tshort/Bradstreet. Students in the Friday afternoon Statistics in Medicine class can choose individual homework assignments based upon these data sets. Further opportunities include writing data set based teaching papers, and potential research topics. In addition, a high level introduction to newer areas of statistical initiatives will be provided.

March 21, 2007

Measuring and Testing Dependence by Correlation of Distances

Prof. Gabor J. Szekely : 3:30 p.m. in SEO 712
Abstract We introduce a simple new measure of dependence between random vectors. Distance covariance (dCov) and distance correlation (dCor) are analogous to product-moment covariance and correlation, but unlike the classical definition of correlation, dCor = 0 characterizes independence for the general case. The empirical dCov and dCor are based on certain Euclidean distances between sample elements rather than sample moments, yet have a compact representation analogous to the classical covariance and correlation. Definitions can be extended to metric-space-valued observations where the random vectors could even be in different metric spaces. Asymptotic properties and applications in testing independence will also be discussed. A new universally consistent test of multivariate independence is developed. Implementation of the test and Monte Carlo results are presented.

April 18, 2007

A Gaussian Calculus for Inference from High Frequency Data

Professor Per Mykland : 3:30 p.m. in SEO 712
Abstract In the econometric literature of high frequency data, it is often assumed that one can carry out inference conditionally on the underlying volatility processes. In other words, conditionally Gaussian systems are considered. This is often referred to as the assumption of ``no leverage effect". This is often a reasonable thing to do, as general estimators and results can often be conjectured from considering the conditionally Gaussian case. The purpose of this paper is to try to give some more structure to the things one can do with the Gaussian assumption. We shall argue in the following that there is a whole treasure chest of tools that can be brought to bear on high frequency data problems in this case. We shall in particular consider approximations involving locally constant volatility processes, and develop a general theory for this approximation. As applications of the theory, we propose an improved estimator of quarticity, an ANOVA for processes with multiple regressors, and an estimator for error bars on the Hayashi-Yoshida estimator of quadratic covariation.

April 20, 2007

Tree-Structured Gatekeeping Procedures

Professor Ajit C. Tamhane : 2 p.m. in SEO 512
Abstract This talk is in two parts. Part I will give a brief introduction to multiple comparison procedures to provide the necessary background for Part II which is the main topic. For brevity and simplicity, we shall restrict to procedures based on p-values only. Parallel and serial gatekeeping procedures have been recently proposed (Westfall and Krishen 2001, and Dmitrienko, Offen and Westfall 2003) for testing hierarchically ordered families of hypotheses. We generalize these procedures to what we call tree-structured gatekeeping procedures. This generalization is necessary to deal with problems involving hierarchically ordered multiple objectives subject to logical restrictions, e.g., in the analysis of multiple endpoints in dose-control studies and in superiority-equivalence testing. The proposed approach is based on the closure principle of Marcus, Peritz and Gabriel (1976) and uses weighted Bonferroni tests for intersection hypotheses. In special cases of parallel or serial tree structures the closed testing procedure can be shown to be equivalent to stepwise procedures, which are easy to implement. Two illustrative clinical trial examples are given. Note: This work is joint with Alex Dmitrienko, Brian Wiens and Xin Wang, and is based on a paper that has recently appeared in Statistics in Medicine.

April 25, 2007

Statistical Inverse Problems in Network Tomography

Professor Vijay Nair : 3:30 p.m. in SEO 712
Abstract The term network tomography characterizes two classes of large-scale inverse problems that arise in the modeling and analysis of computer and communications networks. One class of problems deals with passive tomography where network traffic data are collected at the nodes, and the goal is to reconstruct origin-destination traffic patterns. The second one is active network tomography where the goal is to recover link-level quality of service parameters, such as packet loss rates and delay distributions, from end-to-end path-level measurements. Internet service providers use this to characterize network performance and to monitor service quality. This talk will provide an overview of the network application, the statistical inverse problems that arise, and some recent research in trying to address them with an emphasis on active tomography. This is joint work with George Michailidis, Earl Lawrence, Bowei Xi, and Xiaodong Yang.

April 27, 2007

BIB Designs with Repeated Blocks

Professor Teresa Azinheira Olivira : 2 p.m. in SEO 512
Abstract The Fisher related information of a balanced block design will remain invariant whether or not the design has repeated blocks. This fact can be used theoretically to build a large number of non-isomorphic designs for the same set of design parameters. Further, as it has been shown by UIC researchers and many others later designs with repeated blocks could be used for many different purposes both in experimentations and surveys from finite populations. In this talk the subject will be briefly reviewed and new results on the existence and construction of such designs will be presented. Several unsolved problems for further research will be presented.

May 2, 2007

Investigation on math anxiety among high school students

Hongmei Liu and Nordia Thomas : 3:15 p.m. in SEO 712
Abstract The objective of our project was to evaluate the prevalence of math anxiety for two periods (A and B) of Algebra I students, and to identify if alternative pedagogical strategies help to alleviate math anxiety. From an initial Math Anxiety Scale (MAS) survey we concluded that students did exhibit math anxiety, with most displaying medium to medium-high levels of math anxiety. We conducted a logistic regression of the difference of scores between the Math Anxiety Rating Scale (MARS) survey given before and after a class, and the students' initial math anxiety rating. Using the results of the logistic regression we concluded that the hands-on teaching approach used in Period A helped to alleviate the students' math anxiety more than the traditional lecture approach employed in Period B.

Sept. 5, 2007

A Modeling Approach for Large Spatial Datasets

Prof. Michael Stein : 3:30 p.m. in SEO 636
Abstract For Gaussian spatial processes observed at a large number of irregularly sited locations, exact calculation of the likelihood is generally not possible due to both memory and computational constraints. If we can write the covariance matrix of the observations as a sparse matrix plus a matrix of moderate rank, then both the number of computations and memory requirements can be greatly reduced. The idea is that the sparse term will capture the local behavior of the process and the low rank term the large-scale behavior. This approach is applied to compute likelihood-based estimates of the spatial covariance structure for total column ozone measurements on a global scale. The approach can be judged a success computationally in that likelihoods can be calculated exactly for datasets far too large to carry out the computations for a more general model. However, various diagnostics show problems with the model, so that further work is needed.

Sept. 12, 2007

Graeco-Latin Square Crossover Designs for Higher Order Carryover Effects

Shi Zhao, PhD Candidate : 3:30 p.m. in SEO 712
Abstract Latin square constructed crossover designs balanced for first order carryover effects, commonly referred to as Williams designs, are commonly used in many PK, PD, and other clinical studies. These designs are used to investigate treatment (t) effects in the presence of two nuisance factors, subjects (s) and periods (p), while evaluating and accommodating first order carryover effects with equal precision among treatment comparisons. In some studies, an additional design factor and higher order carryover effects are of interest. For example, in capsaicin cough challenge studies, the additional design factor cough counter (c) is of interest, as is the possibility of higher order carryover effects given the nature of the capsaicin induced cough endpoint, and the human interaction between the cough counters and the study subjects. Graeco-Latin square crossover designs balanced for up to t-1 and c-1 order residual effects need to be constructed. We illustrate for the case where t = p = c = 4. Specifically, two sets of three 4x4 mutually orthogonal Latin square (MOLS) crossover designs with 12 sequences are constructed. Then permutations of each set of three 4x4 MOLS are enumerated. Particular permutations from the first set of three 4x4 MOLS are selected and superimposed upon selected permutations from the second set of three 4x4 MOLS, to construct the desired Graeco-Latin square crossover design with 24 sequences of treatment and evaluator combinations. Using field theory, this result is generalized to any case where t is prime number or a power of a prime number. Open questions related to these results such as: how to partially balance the treatment and evaluator combinations across periods; how to construct designs for numbers of treatments which are not primes or powers of primes; will also be discussed. Appropriate intellectual building blocks will be presented throughout the talk.

Oct. 3, 2007

Polynomial Spline Confidence Bands for Curve Testing

Jing Wang : 3:30 p.m. in SEO 712
Abstract Asymptotically exact and conservative confidence bands are obtained for nonparametric regression function, based on piecewise constant and piecewise linear spline estimation, respectively. Compared to the pointwise nonparametric confidence interval of Huang(2003), the confidence bands are inflated only by a factor of log(n)^{1/2}, similar to the Nadaraya-Watson confidence bands of Hrdle(1989), and the local polynomial bands of Xia(1998) and Claeskens andVan Keilegom(2003). Simulation experiments have provided strong evidence that corroborates with the asymptotic theory. Testing against the linear spline confidence band, the commonly used trigonometric trend is rejected with highly significant evidence for the Leaf Area Index of Aquatic Agriculture land, based on the remote sensing data collected from East Africa.

Oct. 10, 2007

Investigation on voice handicap index

Leping YIn and Wei Zheng : 3:30 p.m. in SEO 712
Abstract The client's objective is to find important factors causing the patients' voice handicap and to investigate the influence of those factors. There are several potential factors of interest: Diagnosis (kind of disease), Gender, Ethnicity, voice therapy (whether the patient complies with doctors ~ R order or not), Age, and Singer or not. All the explanatory factors are categorical variables. We run PROC GLM in SAS and fit an ANOVA model with interactions between the factors. The following factors are found significant in the data analysis: Diagnosis, Singer, Therapy, Gender, and Gender*Diagnosis.

Oct. 17, 2007

Statistical Considerations in Planning and Testing for Multiple Endpoints in Clinical Trials

Dr. Mohammad F. Huque : 3:30 p.m. in SEO 712
Abstract In evaluating that a test treatment is safe and effective in treating a disease, it is often necessary in clinical trials to answer more than one clinically relevant question or to characterize a treatment effect in two or more endpoints. This requires framing clinically relevant multiple hypotheses involving multiple primary and secondary endpoints and treatment comparisons. These multiple hypotheses can be statistically tested according to a strategy that depends on the objectives of the trial and clinical considerations of the disease and the treatment under study. However, such a statistical testing strategy concerning multiple hypotheses involving multiple endpoints is fraught with multiplicity issues, which if ignored can increase the chance of spurious positive findings resulting in false inferences that an effect is shown when there is really no such an effect. In interpreting results from a situation like this, one is more likely to make a false conclusion about the benefit of the test treatment because there are multiple opportunities to choose favorable results from multiple analyses. Therefore, it is necessary that a study protocol of the trial include a clear plan for addressing multiplicity issues and a statistical testing strategy that is appropriate for a given benefit claim of the study treatment. This presentation will examine some basic principles and statistical considerations that can be helpful in better planning and testing of multiple endpoint hypotheses in clinical trials.

Oct. 24, 2007

Semiparametric Analysis of Heterogeneous Data Using Varying-Scale Generalized Linear Models

Prof. Douglas G. Simpson : 3:30 p.m. in SEO 712
Abstract A class of heteroscedastic generalized linear regression models is developed in which a subset of the regression parameters are scaled nonparametrically. Efficient semiparametric inferences are derived for the parametric components of the models. Bootstrap tests for scale heterogenerity are also developed. The models provide an approach to adapt for heterogeneity in the data due factors such as to varying exposures and varying levels of aggregation. The methodology is illustrated with simulations, published data and data from collaborative research on ultrasound safety.

Oct. 31, 2007

Fixed vs. Random Censoring: Is Ignorance Bliss?

Prof. Stephen Portnoy : 3:30 p.m. in SEO 712
Abstract In many situations where censored observations are observed, it is not unreasonable to assume that the censoring values are known for all observations (even the uncensored ones). For example, one of the earliest approaches to ensored regression quantiles was introduced by work of Powell in the mid 1980's. Powell assumed that the censoring values were constant, thus positing observations of the form Y = min(T, c) (where Y is observed and T is the possibly unobserved survival time that is assumed to obey some linear model). More generally, we may be willing to assume that we observe a sample of censoring times {ci} and a sample of censored responses Yi = min(Ti, ci) , a model that could apply to a single sample. In this case, one could use the empirical distributions of the {Yi} and {ci} and take the ratio of empirical survival functions to estimate the survival function of T. This is asymptotically equivalent to applying the Powell method on a single sample. Despite some optimality claims of Newey and Powell, it turns out that the Kaplan Meier estimate is better (asymptotically, and by simulations in finite samples) even though it does not use the full sample of {ci} values. More generally, even in multiple regression settings, the censored regression quantile estimators (Portnoy, JASA, 2003) are better in simulations than Powell's estimator (even for the constant censoring situation for which Powell's estimator was developed). Remarkably, in the one sample case, replacing the empirical function of {ci} by the true survival function (assuming it is known) yields an even less efficient estimator. Thus, it appears that discarding what appears to be pertinent information improves the estimators. The talk will try to quantify and explain this conundrum.

Nov. 7, 2007

Dealing With Missing Values in Clinical Trials From Regulatory Perspectives

Dr. Kooros Mahjoob : 3:30 p.m. in SEO 712
Abstract In some randomized clinical trials, missing values arise due to patients' discontinuation before the end of the trial. As a result, there will be no value/measurement for those patients who dropped out for assessing efficacy at the end of the trial. Such a phenomenon is typical in neurological and psychiatric clinical trials; in fact, in some cases, there are over 40% dropouts. Clearly, analyzing trials data sets that have missing values and then drawing conclusions from the analysis results is a challenging task for FDA statisticians. Dealing with missing values in clinical trials has a long history, which goes back for over two decades. Lots of work, published papers and technical notes have suggested methods to deal with the issue and have talked about the utility of one method over the others. Nevertheless, the reality is that there is no clear-cut solution to the problem. Often, in some trials, missing values is a real problem from the regulatory perspective as to how to make a decision on the drug approval. This presentation will focus on framing the problem, discussing common statistical methods used and the FDA's views on the methods. I will also discuss the result of simulations, bootstrapping using data of actual trials, conducted by FDA colleagues, in comparing the performances of different methods, and the outlines of some new methods.

Nov. 14, 2007

Consistent long-memory parameter estimation in a LARCH time series model and its connection to the Hurst parameter of the fractional Brownian motion

Michael Levine : 3:30 p.m. in SEO 712
Abstract We investigate several possible strategies for consistently estimating the so-called Hurst parameter H responsible for the long-memory property in a special class of nonlinear ARCH-type models popularly known as LARCH, as well as in the continuous-time Gaussian stochastic process named fractional Brownian motion (fBm). Several estimation methods are discussed, including a conditional MLE method and a local Whittle-type estimation procedure. The conditional MLE is proved to be consistent and a Portmanteau-type test for model validation is established. By constructing the LARCH and fBm processes on a common probability space, and showing the convergence of various partial sums of the former to the latter in mean squared, we can propose a specially designed conditional maximum likelihood method for estimating the fBm's Hurst parameter. In keeping with the popular financial interpretation of ARCH-type models, all estimators are based only on observation of the "returns" of the model and not on the "volatilities".

Nov. 16, 2007

Statistical Investigation on High Hospital Cost of Hispanic Patients

Yuan Xu : 2 p.m. in SEO 712
Abstract Patients with heart disease usually stay in the hospital for certain days before they are cured and then leave the hospital. For each patient, the amount of money that he/she spends in the hospital usually is different. Some people pay more, some the other people pay less. Now we have total 143211 records of such patients. For each patient's record, we have his/her total expenditure during stay in hospital and a lot of other information such as his/her sex, age, race, length of stay in hospital, different disease history for example once having shock or not, different hospital conditions, different financial conditions for example insurance and so on. It is observed that in average Hispanic Women or Hispanic Men tend to spend more money during their stay in hospital than other race group of people. Our objective is to give a good explanation on what is the cause for the above phenomenon. After discussion with medical doctors who provided valuable suggestions, we started with 25 suspected variables. After running the linear regressions for the total and each of the 8 sub-groups of patients, we excluded many of these variables and concentrated our investigation on 5 left variables. We applied different kinds of statistical tests on these 5 variables and found that most likely only two of them may contribute to the higher cost of Hispanic patients. By medical doctor's opinion, we excluded one of these two variables and kept the last one as our investigation result. We claim that the special behavior of Hispanic patients regarding to this variable actually causes them to pay more in the hospital.

Nov. 28, 2007

Construction of Simultaneous Confidence Bands in Time Series

Weibiao Wu : 3:30 p.m. in SEO 712
Abstract I will talk about statistical inference of trends in mean non-stationary models, and mean regression and conditional variance (or volatility) functions in nonlinear stochastic regression models. Simultaneous confidence bands are constructed and the coverage probabilities are shown to be asymptotically correct. The Simultaneous confidence bands are useful for model specification problems in nonlinear time series. The results are applied to environmental and financial data-sets.

Jan. 23, 2008

Estimation of vaccine efficacy and the vaccination threshold

Prof. Qizhi Chen : 3:30 p.m. in SEO 712
Abstract This paper considers the effect of imperfect vaccination in a susceptible infected removal (SIR) epidemic model. The minimum proportion of the population that needs to be vaccinated to prevent a major epidemic depends on the vaccine efficacy and the basic reproductive rate for the SIR model, allowing for imperfect and variable vaccination. Martingale theory is used to derive estimates and associated standard errors for these parameters. Asymptotic properties of the resulting estimators are investigated. Data for a mumps outbreak are used as an illustrative example.

Jan. 29, 2008

On Large Margin Hierarchical Classification

Junhui Wang : 3 p.m. in SEO 636
Abstract Hierarchical classification is critical to knowledge and context management as well as knowledge exploration, as in gene function classification and discovery and document categorization. In hierarchical classification, an input is classified by a structured hierarchy. In a situation as such, the central issue is how to effectively utilize inter-class relationship to improve the generalization performance of flat classification ignoring such dependency. In this talk, a novel large margin method based on constraints characterizing multi-path hierarchy is presented within the framework of regularization. In particular, I will discuss three aspects: (1) the idea and methodology development; (2) computational tools; (3) a statistical learning theory. Numerical examples will be provided to demonstrate the advantage of our proposed methodology against other existing competitors. An application to gene function prediction and discovery will be discussed.

Feb. 28, 2008

Power analysis for the Zenker's diverticulum trial

Yufei Chen and Weiyun Zheng : 1 p.m. in SEO 712
Abstract Zenker's diverticulum is a diagnosis that primarily affects individuals in the seventh and eighth decades. The current mainstay of treatment is surgical. Due to the advanced age of the population, many are not prime surgical candidates and have a higher probability of sequelae from general anesthetics. A newer method of delivery is injections that are given in clinic under EMG guidance. This allows patients who have comorbidities that make them poor surgical candidates the ability to receive treatments. To see the value of the botulium toxin injections, our study will look at the subjective symptomatic improvement noted by patients after receiving the injections. Patients who will receive botox injections in the clinic setting are asked to fill out a survey about their current status before the injections and the improvements they will notice after treatment in eating ability, normalcy of diet, and understandability of speech. With the distribution of patients populations to each category, we firstly use independent Z-test to get a rough total sample number with sufficient power. Second, we use advanced model with conditional probability transition matrix to simulate the whole process. The two results will be compared and the advantage of the later will be discussed.

March 5, 2008

Experimental Design for Nonlinear Mixed-Effects Models with Application to Studies of HIV Dynamics

Dr. Cong Han : 3:30 p.m. in SEO 712
Abstract This presentation will review design issues for studies of HIV dynamics that use a nonlinear mixed-effects model. A method based on a first-order approximation, used for similar design issues in pharmacokinetics and pharmacodynamics, will be reviewed and discussed. Limitations of this method will be discussed and other methods will be described, including methods based on the exact calculation of the Fisher information matrix and Bayesian methods, which, although computationally intensive, provide alternatives.

March 12, 2008

Asymptotic properties in ARCH(p)-time series

Prof. Fuxia Cheng : 3:30 p.m. in SEO 712
Abstract ARCH(p)-model has found much interest in financial econometrics. It was introduced by Engle(1982) in order to provide a framework in which so-called volatility clusters may occur, i.e., periods of high and low (conditional) variances depending on past values of the series. The model was later extended into various directions. In most of the work, the main focus has been on estimating the unknown parameters. But it is of interest and of practical importance to know the nature of the innovation distribution. Actually, if the distribution of the innovation is unspecified, the parametric component only partly determines the distribution behavior. It is as important to investigate the distribution of the innovation as estimating the parameters. In this talk, we consider the consistency and the asymptotic distribution of the innovation density estimators in ARCH(p)-time series. We also extend the central limit theorem (CLT) and the strong law of large number (SLLN) to the average of the residuals.

March 19, 2008

Semiparametric detection of significant activation for brain fMRI

Prof. Chunming Zhang : 3:30 p.m. in SEO 712
Abstract Functional magnetic resonance imaging (fMRI) aims to locate activated regions in human brains when specific tasks are performed. The conventional tool for analyzing fMRI data applies some variant of the linear model, which is restrictive in modeling assumptions. To yield more accurate prediction of the time-course behavior of neuronal responses, the semiparametric inference for the underlying hemodynamic response function is developed to identify significantly activated voxels. Under mild regularity conditions, we demonstrate that a class of the proposed semiparametric test statistics, based on the local linear estimation technique, follow chi-squared distributions under null hypotheses for a number of useful hypotheses. The asymptotic power functions of the constructed tests are derived under the fixed and contiguous alternatives. Furthermore, a new false discovery rate approach which incorporates spatial information of voxel-wise p-values is devised for detecting the regions of activation. Simulation evaluations and real fMRI data application suggest that the semiparametric inference procedure provides more efficient detection of activated brain areas than the popular imaging analysis tools.

April 2, 2008

Entropy, folding, and function of biomolecules through Monte Carlo sampling

Prof. Jie Liang : 3:30 p.m. in SEO 712
Abstract The three dimensional structures of biomolecules such as proteins and RNAs are the basis of their biological functions. For RNA molecule, conformational entropy is important for stability and folding. However, it is challenging to either measure or compute conformational entropy associated with long loops. We develop optimized discrete $k$-state models of RNA backbone and estimate entropy of hairpin, bulge, internal loop, and multibranch loop of long length using an efficient sequential Monte Carlo sampling method. The estimated entropy indicate that the Jacobson-Stockmayer model has large errors for bulge, internal, and multibranch loops. For protein, we study the transition state ensemble. By generating effective samples under various experimentally derived constraints, we characterize the transition state ensemble (TSE) during protein folding. As TSE is short-lived, the size and shape of conformations in TSE have been elusive. For the protein acylphosphatase, we found TSE has diverse conformations. In contrast to previous results, we found overall TSE can be very different from native structure of proteins (with RMSD>12A). To predict protein functions, we develop a method by matching local surfaces based on estimated evolutionary information specific to individual binding region via a Bayesian Monte Carlo approach using a continuous-time Markov model. Our method provides a probabilistic model which characterizes protein binding activities that may involve multiple substrates or ligands. (Joint work with Rong Chen, Ming Lin, Zheng Ouyang, Jeffrey Tseng, and Jian Zhang) (please visit http://www.uic.edu/~jliang for further information).

April 9, 2008

Adaptation of Clinical Trial Design

H.M. James Hung, PhD : 4 p.m. in SEO 636
Abstract A topic of great interest in the recent decade is design adaptation for clinical trials using the data accumulating during the course of the trial. There are good reasons for such modification of design features. Design adaptations include sample size re-estimation, enrichment of patient population, dropping a treatment arm, etc. In this presentation I shall give an introduction of this topic and a brief overview of critiques and comments to such adaptation in the literature.

April 16, 2008

An Application of Path Analysis in the Design of Clinical Trials

Dr. Yili Pritchett : 3:30 p.m. in SEO 712
Abstract Path analysis, first developed in the field of genetics and actively used in sociology, refers to a modeling approach for causal relationships (Wright, 1934). In the setting of clinical research, path analysis can be a useful tool to demonstrate an independent treatment effect on a disease state, which might have causal relationships with other disease states that can be treated by the same therapy. In this presentation, the idea of using path analysis in clinical trial design will be introduced. In this approach, pre-specified causal relationships can be modeled by structural equations so that the treatment effect on the disease state of interest (direct effect) will be tested after accounting for the treatment effects on the other conditions (indirect effects). The total treatment effect can be decomposed as the sum of the direct and the indirect effects, and the statistical significance of each effect can be tested. The idea and the approach will be illustrated by concrete examples where an antidepressant was investigated for its effect on the management of different types of pain. Further statistical discussions will be given on the topics of the invariance between ordinary linear regression and standardized linear regression, and the generalization from ordinary linear model to generalized linear model.

April 23, 2008

Parameter Estimation and Model Selection in Graphical Models

Prof. Xin Gao : 3:30 p.m. in SEO 712
Abstract The recent years have witnessed the increasing interest in the study of graphical models. In this talk, I will discuss two related parameter estimation and model selection problems in graphical models. The first problem is to estimate the concentration matrix of a Gaussian graphical model. We propose to estimate the concentration matrix using the penalized likelihood method with the smoothly clipped absolute deviation (SCAD) penalty. The method leads to a sparse and shrinkage estimator of the concentration matrix. Using proper choice of the regularization parameter, the proposed method automatically and consistently selects the true graphical structure and produces estimator that is as efficient as the oracle estimator. We further establish the consistency of the BIC criterion to identify the true graphical structure when used with the SCAD penalty function. The second problem is regarding the graphic model with multivariate hidden Markov structure. For such high-dimensional data with complicated dependency structure, we propose to use composite likelihood approach and especially we develop COMP-EM algorithm to perform the parameter estimation in the presence of incomplete data. The composite likelihood based information criterion was employed to select the best network structure.

April 24, 2008

Intervention Analysis in a Water Usage Study

Matthew J. Bourque and Xuejing Wang : 2:15 p.m. in SEO 712
Abstract A water main bringing water into a Chicago suburb was damaged on or about January 5, 2002. The damage was not discovered until February 11, 2004, for a total of approximately 765 days. We have daily water usage data from 1989 through 2007, and want to determine (i) whether there is a statistically significant increase in water flow during the period from January 2002 until February 2004 and (ii) if there is a significant increase, to estimate the amount of extra water flow during that time. We use intervention analysis to fit an ARIMA model, and from the estimates of the parameters we get the result. We will also discuss more generally about intervention analysis.

April 30, 2008

Lasso and Shrinkage Estimation in Generalized Linear Models

Prof. Ejaz Ahmed : 3:30 p.m. in SEO 712
Abstract We consider the estimation problem for the parameters of generalized linear models which may have a large collection of potential predictor variables and some of them may not have influence on the response of interest. In this situation, selecting the statistical model is always a challenging problem. In the context of two competing models, we demonstrate the relative performances of shrinkage and classical estimators based on the asymptotic analysis of quadratic risk functions. We demonstrate that the shrinkage estimator outperforms the maximum likelihood estimator uniformly. For comparison purpose, we also consider the Park and Haste type estimator (variant of lasso estimator) for generalized linear models. This comparison shows that shrinkage method performs better than the lasso type estimation method when the dimension of the restricted parameter space is large. This talk ends with real-life example showing the value of new method in practice. More, specifically, we consider South African heart disease data, which was collected on males in a heart disease high-risk region of Western Cape, South Africa.

May 1, 2008

Modeling and Analyzing High-Frequency Financial Data

Prof. Yazhen Wang : 3 p.m. in SEO 636
Abstract Volatilities of asset returns are central to the theory and practice of asset pricing, portfolio allocation, and risk management. In financial economics, there is extensive research on modeling and forecasting volatility up to the daily level based on Black-Scholes, diffusion, GARCH, stochastic volatility models and implied volatilities from option prices. Nowadays, thanks to technological innovations, high-frequency financial data are available for a host of different financial instruments on markets of all locations and at scales like individual bids to buy and sell, and the full distribution of such bids. The availability of high-frequency data stimulates an upsurge interest in statistical research on better estimation of volatility. This talk will start with a review on low-frequency financial time series and high-frequency financial data. Then I will introduce popular realized volatility computed from high-frequency financial data and present my work on multi-scale methods for analyzing jump and volatility variations and matrix factor models for handling large size volatility matrices.

Sept. 24, 2008

Modeling Trade Direction

Prof. Dale Rosenthal : 4:15 p.m. in SEO 612
Abstract The problem of classifying trades as buys or sells is examined. I propose estimated quotes for midpoint and bid/ask tests and a modeling approach to classification. Prevailing quotes are estimated using flexible approximations to the distribution for delays of quotes relative to trade timestamps. Classification is done by a generalized linear model which includes improved versions of midpoint, tick, and bid/ask tests. The model also considers the relative strengths of these tests, can account for market microstructure peculiarities, and allows for autocorrelations and cross-correlations in trade direction. The correlation modeling corrects for pseudoreplication, yielding more accurate standard errors and fixed effect estimates. Further, the model estimates probabilities of correct classification. The model is compared to various trade classification methods using a sample of 2,836 domestic US stocks from an unexplored, recent, and readily-available data set. Out of sample, modeled classifications are 1-2% more accurate overall than current methods; this improvement is consistent across dates, sectors, and locations relative to the inside quote. For Nasdaq and NYSE stocks, 1% and 1.3% of the improvement comes from using relative strengths of the various tests; 0.9% and 0.7% of the improvement, respectively, comes from using some form of estimated quotes. For AMEX stocks, a 0.4% improvement is attributed to using a lagged version of the bid/ask test. I also find indications of short- and ultra-short-term alpha.

Oct. 1, 2008

Hypothesis tests in algebraic statistical models

Prof. Mathias Drton : 4:15 p.m. in SEO 612
Abstract Many statistical models are defined in terms of polynomial constraints, or in terms of polynomial or rational parametrizations. Such algebraic models include, for instance, factor analysis and instrumental variable models, latent class models, and more generally, discrete and Gaussian graphical models with hidden variables. Statistical inference in hidden variable models is complicated by the fact that the models' parameter spaces are typically not smooth. This is the motivation for this talk that considers testing a null hypothesis with singularities in algebraic models. The focus will be on the large-sample asymptotic behavior of likelihood ratio and Wald tests.

Oct. 8, 2008

Probability Estimation for Large Margin Classifiers

Prof. Junhui Wang : 4:15 p.m. in SEO 612
Abstract Large margin classifiers have proven to be effective in delivering high predictive accuracy, particularly those focusing on the decision boundaries and bypassing the requirement of estimating the class probability given input for discrimination. As a result, these classifiers may not directly yield an estimated class probability, which is of interest itself in many real applications. In this talk, I will present a novel method to estimate the class probability through sequential weighted classifications, by utilizing features of interval estimation of large margin classifiers. In particular, I will discuss four aspects: (1) the idea and methodology development; (2) tuning parameter selection; (3) regularization solution path; (4) a statistical learning theory. Numerical examples will be provided to demonstrate the advantage of our proposed methodology.

Oct. 15, 2008

On the Impact of Formative Assessment on Student Motivation, Achievement, and Conceptual Change

Prof. Yue Yin : 4 p.m. in SEO 612
Abstract Formative assessment was hypothesized to have a beneficial impact on students' science achievement and conceptual change, either directly or indirectly by enhancing motivation. We designed and embedded formatives assessments within an inquiry science unit. Twelve middle-school science teachers with their students were randomly assigned either to an experimental group (N = 6), provided with embedded formative assessment, or control group (N = 6). Teachers varied significantly as to their impact on student motivation, achievement, and conceptual change. But the impact of the formative assessment treatment on these outcomes was not statistically significant. Variation in both teachers' classroom management and the degree to which they used informal formative assessment, regardless of group, were conjectured as possible reasons for the absence of an overall formative assessment effect.

Oct. 22, 2008

Crossover Designs under Subject Dropouts

Shi Zhao, Ph. D. : 4:15 p.m. in SEO 612
Abstract Crossover experiments are used for comparing the responses to various different stimuli or treatments in areas ranging from psychology and human factor engineering to medical and agricultural applications. They are widely used in the pharmaceutical industry. There is an extensive literature that assures us that a carefully designed crossover study will produce a wealth of information that will enable inference with high precision. This is based on the implicit, but critical, assumption that the experiment will yield all the planned observations. Yet in many situations, such as clinical trials, there is a substantial probability that some subjects will drop out of the study prior to the completion of their treatment sequence. Low, Lewis and Prescott (1999) observed that a dropout rate of between 5% and 10% is not uncommon and, in some areas, can be as high as 25%. They gave an example of a design in four periods based on a Williams Latin square where there is substantial loss of information if some observations are unavailable in period 4. Indeed, if all observations in the final period are not available, the design becomes disconnected, i.e., elementary contrasts are no longer all estimable. Majumdar, Dean and Lewis (2005) studied the maximum loss in uniformly balanced repeated measurements designs (UBRMDs) in t periods when subjects may drop out after period t-m. We will further study UBRMDs under the subject dropouts. (1) We will derive the "best" UBRMDs for the situation where all subjects may drop out in the final period and provide methods for constructing these designs. (2) We will study UBRMDs under subject dropout for the model where subject effects are random and show that the Low, Lewis and Prescott (1999) result on lack of connectedness of the Williams Latin Square of order 4 is no longer valid. Compound symmetry and AR (1) covariance structures will be considered. (3) Expected loss under various dropout probabilities will be studied.

Oct. 29, 2008

On False Discovery Control under Dependence -- (Cancelled)

Prof. Wei Biao Wu : 4:15 p.m. in SEO 612
Abstract A popular framework for false discovery control is the random effects model in which the null hypotheses are assumed to be independent. I will generalize this random effects model to a conditional dependence model which allows dependence between null hypotheses. The dependence can be useful to characterize the spatial structure of the null hypotheses. Asymptotic properties of false discovery proportions and numbers of rejected hypotheses are explored and a large-sample distributional theory is obtained. The talk is based on the paper: Wu, W. B. (2008) On false discovery control under dependence, Ann. Statist. 36, 364--380.

Nov. 5, 2008

Consequences of research hypotheses captured by special types of independence graph

Prof. Nanny Wermuth : 4:15 p.m. in SEO 612
Abstract A joint density of several variables may satisfy a possibly large set of independence statements, called its independence structure. Often this structure is fully representable by a graph that consists of nodes representing variables and of edges that couple node pairs. We consider joint densities of this type, generated by a stepwise process in which all variables and dependences of interest are included. Otherwise, there are no constraints on the type of variables or on the form of the distribution generated. For densities that then result after marginalising and conditioning, we derive what we name the summary graph. It is seen to capture precisely the independence structure implied by the generating process, it identifies dependences which remain undistorted due to direct or indirect confounding and it alerts to possibly severe distortions of these two types in other parametrizations. We use operators for matrix representations of graphs to derive matrix results and translate these into to special types of path.

Nov. 12, 2008

How fast does a Brownian motion move on a manifold?

Prof. Elton P Hsu : 4:15 p.m. in SEO 612
Abstract How fast does a transient Brownian motion escape to infinity is an interesting question. For example, it is well known that Brownian motion in a euclidean space of dimension 3 and higher escapes to infinity at approximately the rate of the square root of time. Which geometric quantity controls the rate of escape is an interesting question. We will explain that the speed of Brownian motion can be effectively controlled by the volume growth rate of a complete Riemannian manifold. A precise integral criterion is obtained which in many cases is sharp.

Nov. 19, 2008

Large-Scale Prediction Problems

Prof. Brad Efron : 4:15 p.m. in Lecture Center D2
Abstract Classical prediction methods such as Fisher's linear discriminant function were designed for small-scale problems, where the number N of candidate predictors was much smaller than the number of observations n. Modern scientific devices often reverse this situation. A micro- array analysis, for example, might include n=100 subjects measured on N=10,000 genes, each of which is a potential predictor. I will discuss "Ebay", an empirical Bayes prediction algorithm designed to handle N >> n situations. It is closely related to the Shrunken Centroids algorithm of Tibshirani, Hastie, Narasimhan, and Chu.

Dec. 3, 2008

Cancelled

Prof. Min Zhang : 4:15 p.m. in SEO 612
Abstract High throughput biotechnologies such as microarray and next-generation sequencing permit simultaneous measurements of enormous bodies of expression and sequence information. However, the number of biological samples is much smaller compared to the number of available predictors. Statistically, we are challenged by the large number of parameters but small number of observations. To tackle this issue, we proposed a two-step variable selection procedure to reduce the dimension in the first stage where Gibbs sampler was developed to stochastically search through low-dimensional subspaces. With reduced number of variables, either Bayesian variable selection or traditional approaches can be employed in the second stage. The methods are evaluated via simulation studies and we also applied them to real data sets, including QTL mapping data and gene expression data.

Jan. 21, 2009

Portmanteau Tests in Time Series

Prof. Xiaofeng Shao : 4:15 p.m. in SEO 612
Abstract This talk consists of two parts. In the first part, we will talk about testing for white noise and its applications to goodness-of-fit of long memory time series models. The limitation of the current asymptotic theory for portmanteau tests will be pointed out and new theoretical results will be discussed. In the second part, we will introduce generalized portmanteau type test statistics in the frequency domain to test independence between two stationary time series. Unlike the existing tests, each time series is allowed to possess short memory, long memory or anti-persistence. Under the null hypothesis of independence, the asymptotic null distributions of the proposed statistics are standard normal. The results from a simulation study will also be presented.

Jan. 28, 2009

On False Discovery Control under Dependence

Prof. Wei Biao Wu : 4:15 p.m. in SEO 612
Abstract A popular framework for false discovery control is the random effects model in which the null hypotheses are assumed to be independent. I will generalize this random effects model to a conditional dependence model which allows dependence between null hypotheses. The dependence can be useful to characterize the spatial structure of the null hypotheses. Asymptotic properties of false discovery proportions and numbers of rejected hypotheses are explored and a large-sample distributional theory is obtained. The talk is based on the paper: Wu, W. B. (2008) On false discovery control under dependence, Ann. Statist. 36, 364--380.

Feb. 4, 2009

A Bayesian Nonparametric Causal Model For Observational Studies

Prof. George Karabatsos : 3 p.m. in SEO 636
Abstract Often, causal inference is conducted on the basis of the randomized experiment. However, in many settings, a randomized experiment is infeasible, because treatments cannot be directly assigned to subjects. Specifically, it may not be possible for the investigator to assign treatments to subjects, because of ethical concerns, or because of excessive expense in terms of time or money. In such settings, causal inference needs to be undertaken in an observational study, where the subjects received different treatments, but the investigator did not assign the treatments, and therefore the treatment assignment probabilities are unknown. Typically, in the practice of causal inference from observational studies, a parametric model is assumed for the joint population density of potential outcomes and treatment assignments, and possibly this is accompanied by the assumption of no hidden bias. However, both assumptions are questionable for real data, the accuracy of causal inference is compromised when the data violates either assumption, and the parametric assumption precludes capturing a more general range of density shapes (e.g., heavier tail behavior and possible multi-modalities in the joint density). We introduce a flexible, Bayesian nonparametric causal model to provide more accurate causal inferences. The model makes use of a stick-breaking prior distribution, which has the flexibility to capture any multi-modalities, skewness and heavier tail behavior in this joint population density, while accounting for hidden bias. We prove the asymptotic consistency of the posterior distribution of the model. Also, we illustrate our Bayesian nonparametric causal model through the analysis of small genetic data set, and a large data set of Chicago public schools.

Feb. 11, 2009

Multi-objective Optimal Experimental Designs in Event-Related fMRI Studies

Prof. Abhyuday Mandal : 4:15 p.m. in SEO 612
Abstract Functional magnetic resonance imaging (fMRI) is considered one of the leading technologies for studying human brain activity in response to mental stimuli. With sophisticated allocations of stimuli, researchers can gather valuable fMRI time series and acquire precise information about human brain activity. However, due to the nature of fMRI experiments, the underlying design space is very large and irregular. This makes it difficult to find an optimal design that simultaneously accomplishes various goals of a study and fulfills the scientific restrictions. Here we propose an efficient approach to find optimal experimental designs for event-related functional magnetic resonance imaging (ER-fMRI). We consider multiple objectives, including estimating the hemodynamic response function (HRF), detecting activation, circumventing psychological confounds and fulfilling customized requirements. Taking into account these goals, we formulate a family of multi-objective design criteria and develop a genetic-algorithm-based technique to search for optimal designs. Our proposed technique incorporates existing knowledge about the performance of fMRI designs, and its usefulness is shown through simulations. We also find designs yielding higher estimation efficiencies than m-sequences. When the underlying model is with white noise and a constant nuisance parameter, the stimulus frequencies of the designs we obtained are in good agreement with the optimal stimulus frequencies derived by Liu and Frank, 2004, NeuroImage 21, 387-400. In terms of CPU time and achieved design efficiency, we demonstrate that our approach outperforms the methodologies known hitherto. (Joint research with Ming-Hung (Jason) Kao, John Stufken and Nicole Lazar)

Feb. 13, 2009

Can Statistics Help Improve US Elections?

Dr. Arlene Ash : 2 p.m. in SEO 612
Abstract US elections are very complicated, providing many opportunities for inadvertent and malicious errors. I will examine the statistical evidence regarding a "failed election" in Florida's 13th Congressional District in 2006 and discuss some lessons learned. I will discuss additional ways in which elections can be problematic and how statistics can be used to identify problems and explore solutions. Finally, I will describe work in progress with state election officials to improve the efficiency and effectiveness of post-election audits.

Feb. 25, 2009

An invitation to algebraic statistics

Sonja Petrovic : 4:15 p.m. in SEO 612
Abstract Algebraic statistics is a maturing discipline whose main focus is the study of statistical models using the tools from algebraic geometry and computational algebra. The main concept is that statistical models are algebraic varieties. Algebraic approach can be used to provide Markov bases for the models, or to compute the maximum likelihood degree. Some of the best studied models so far are contingency tables, conditional independence models and graphical models including latent class. Algebraic statistics has also found applications in computational biology and phylogenetics. Some recent work shows that it can be used as a powerful tool for phylogenetic tree reconstruction and for model identifiability problems. This will be an introductory talk to explain some of the main concepts of the field, illustrated on a few examples.

March 4, 2009

A New Method to Compare the Haplotype Distributions between Populations

Prof. Liping Tong : 3 p.m. in SEO 636
Abstract Accurate characterization of haplotype structure and diversity is a key challenge in statistical genetics. Attempts to apply findings from genome wide association studies to populations not included in the discovery phase present unique challenges in terms of the statistical methods. In this talk, I propose a new statistic to assess and compare the haplotype variations among populations which is particularly suited to this emerging challenge. Subsequently I show that this statistic follow a weighted chi-square distribution and how to use a chi-square distribution to approximate it. This approximation is very important since no other haplotype similarity tests have (correctly) used approximate theoretical distributions. In stead, the computational intensive permutation tests are generally performed, which limit the application of haplotype-based comparisons to the whole genome wide studies. In the simulation studies, I first discuss the performance of the approximate distribution under different definitions of similarity matrix, and then compare the power of my new method with the ones proposed by others. At last, this method is applied to the HapMap data to test population differences based on haplotypes on chromosome 2 in the region surrounding the LCT gene (135.3-136.9 Mb).

March 18, 2009

Asymptotics of the Theil-Sen estimators in a multiple linear regression model when covariates are deterministic

Prof. Fang Li : 4 p.m. in SEO 612
Abstract In this talk, we introduce the Multivariate Theil-Sen Estimators (MTSE) in a multiple linear regression model generalizing the TSE in a simple linear model. We demonstrate that the proposed estimators are robust with bounded influence function. We also show that the MTSE is consistent under very mild conditions. When the covariates are independent and identically distributed, the related criterion statistics defining the MTSE are the standard U-statistics. However when the covariates are deterministic, the related criterion statistics are no longer U-statistics. We then use the Convexity Lemma of Pollard to prove its asymptotic normality under mild condition. This method can also be easily generalized to the case when the covariates are independent and identically distributed.

April 8, 2009

Nonparametric Smoothing With Infinite Order Kernels

Prof. Tim McMurry : 4:15 p.m. in SEO 612
Abstract I will discuss the usefulness of infinite order kernels in nonparametric smoothing problems. Particular focus will be paid to nonparametric regression, but the ideas apply to many related problems, including density and spectral density estimation. It will be shown that estimators using infinite order kernels are completely and automatically adaptive to the underlying function being estimated, produce favorable results in simulation, and substantially facilitate construction of confidence intervals.

April 15, 2009

Semiparametric Regression Models With Measurement Errors

Prof. Hua Liang : 4:15 p.m. in SEO 636
Abstract We investigated two semi-parametric models, partially linear model and generalized partially linear model, with error-prone covariates. A correction-for-attenuation method was developed for estimating the parameter of interest in the partially linear model. The resulting estimator was shown to be consistent and its asymptotic distribution theory has been derived. Consistent standard error estimates using sandwich-type ideas were also developed. For generalized partially linear model, we proposed estimators of parameter and nonparametric function by using local linear regression, simulation extrapolation technique, and generalized estimating equation. The asymptotic normality of the estimators of the parameter, the bias and variance of the estimators of the nonparametric component were derived under appropriate assumptions. We illustrated the numerical performance of the proposed methods via simulation and examples, discussed the potential topics for further work.

April 22, 2009

Bayesian variable selection for high dimensional models with applications in genomics

Prof. Min Zhang : 4:15 p.m. in SEO 612
Abstract High throughput biotechnologies such as microarray and next-generation sequencing permit simultaneous measurements of enormous bodies of expression and sequence information. However, the number of biological samples is much smaller compared to the number of available predictors. Statistically, we are challenged by the large number of parameters but small number of observations. To tackle this issue, we proposed a two-step variable selection procedure to reduce the dimension in the first stage where Gibbs sampler was developed to stochastically search through low-dimensional subspaces. With reduced number of variables, either Bayesian variable selection or traditional approaches can be employed in the second stage. The methods are evaluated via simulation studies and we also applied them to real data sets, including QTL mapping data and gene expression data.

April 29, 2009

Crossover Designs under Random Subject Effects

Wei Zheng, PhD candidate : 4:15 p.m. in SEO 612
Abstract The statistical optimality and efficiency of crossover designs for the purpose of comparing several test treatments with a control treatment when the subject effects are random depend heavily on the unknown ratio theta of the variance of subject effects and the error variance. However, it is proved that if the class of competing designs contains a totally balanced test-control incomplete crossover designs (TBTCI), as defined by Hedayat and Yang (2005), then this TBTCI design is simultaneously A- and MV-optimal for all values of theta. This result is essentially a generalization of a result in Hedayat and Yang (2005) since their statistical model is based on fixed subject effects, where the Fisher information matrix would be identical to that of random subject effect model when theta goes to infinity. Partial works on the construction of the designs are carried out.

Sept. 2, 2009

Consistent variable selection in additive models

Prof. Lan Xue : 3 p.m. in SEO 636
Abstract We propose a penalized polynomial spline method for simultaneous model estimation and variable selection in additive models. It approximates nonparametric functions by polynomial splines, and minimizes the sum of squared errors subject to an additive penalty on norms of spline functions. This approach sets estimators of certain function components to zero, thus performing variable selection. Under mild conditions, we show that the newly proposed method estimates the non-zero function components in the model with the same optimal mean square convergence rate as the standard polynomial spline estimators, and correctly sets the zero function components to zero with probability approaching one, as $n$ goes to infinity. Besides being theoretically justified, the proposed method is easy to understand and straightforward to implement. Extensive Monte Carlo simulation studies show the newly proposed method compares favorably with the existing ones in finite sample performance. We also illustrate the use of the proposed method by analyzing two data sets.

Sept. 9, 2009

Penalized orthogonal-components regression for large p small n data

Prof. Dabao Zhang : 3 p.m. in SEO 636
Abstract We propose a penalized orthogonal-components regression (POCRE) for large p small n data. Orthogonal components are sequentially constructed to maximize, upon standardization, their correlation to the re- sponse residuals. A new penalization framework, implemented via empiri- cal Bayes thresholding, is presented to effectively identify sparse predictors of each component. POCRE is computationally efficient owing to its se- quential construction of leading sparse principal components. In addition, such construction offers other properties such as grouping highly correlated predictors and allowing for collinear or nearly collinear predictors. With multivariate responses, POCRE can construct common components and thus build up latent-variable models for large p small n data. This is an joint work with Yanzhu Lin and Min Zhang.

Sept. 16, 2009

An Application of The Implicit Function Theorem in Strong Consistency of Parameter Estimators

Prof. Ahmad Reza Soltani : 3 p.m. in SEO 636
Abstract An approach for proving the strong consistency of certain estimators for unknown parameters in the context of statistical inference is given. This approach is based on an application of the Implicit Function Theorem in Hilbert spaces, and can be applied to the random samples consisting of univariate, multivariate or infinite dimensional random elements.

Sept. 23, 2009

Bootstrap Consistency for General Semiparametric M-estimation

Prof. Guang Cheng : 3 p.m. in SEO 636
Abstract Consider M-estimation in a semiparametric model that is characterized by a Euclidean parameter of interest and a nuisance function parameter. We show that, under general conditions, the bootstrap is asymptotically consistent in estimating the distribution of the M-estimate of Euclidean parameter; this is, the bootstrap distribution asymptotically imitates the distribution of the M-estimate. We also show that the bootstrap confidence set has the asymptotically correct coverage probability. These general conclusions hold, in particular, when the nuisance parameter is not estimable at root-n rate. Our results provide a theoretical justification for the use of bootstrap as an inference tool in semiparametric modelling and apply to a broad class of bootstrap methods with exchangeable bootstrap weights. A by-product of our theoretical development is the second order asymptotic linear expansion of the (bootstrap) M-estimate. Joint work with Jianhua Huang at Texas A&M University.

Sept. 30, 2009

Quotient Correlation: A New Light of Measuring Variable Associations and Testing Hypotheses of Independence and Tail Independence

Prof. Zhengjun Zhang : 3 p.m. in SEO 636
Abstract Various correlation measures have been introduced in statistical inferences and applications. Each of them may be used in measuring association strength of the relationship, or testing independence, between two random variables. \textsl{The quotient correlation} is defined here as an alternative to Pearson's correlation that is more intuitive and flexible in cases where the tail behavior of data is important. It measures nonlinear dependence where the regular correlation coefficient is generally not applicable. One of its most useful features is a test statistic that has high power when testing nonlinear dependence in cases where the Fisher's $Z$-transformation test may fail to reach a right conclusion. Unlike most asymptotic test statistics, which are either normal or $\chi2$, this test statistic has a limiting gamma distribution (henceforth \textsl{the gamma test statistic}). More than the common usages of correlation, the quotient correlation can easily and intuitively be adjusted to values at tails. This adjustment generates two new concepts -- the tail quotient correlation and the tail independence test statistics, which are also gamma statistics. Due to the fact that there is no analogue of the correlation coefficient in extreme value theory, and there does not exist an efficient tail independence test statistic, these two new concepts may open up a new field of study. In addition, an alternative to Spearman's rank correlation: a rank based quotient correlation is also defined. The advantages of using these new concepts are illustrated with simulated data, and real data analysis of internet traffic, tobacco markets, financial markets...

Oct. 7, 2009

Particle Methods for General Mixtures

Prof. Hedibert Lopes : 3 p.m. in SEO 636
Abstract This paper develops efficient sequential learning methods for the estimation of general mixture models, by working directly with particles based on conditional sufficient information. We provide an alternative to existing inference techniques that will be especially relevant in on-line estimation settings and for large,high-dimensional data-sets. With each new observation, particles are updated in two steps: first, resampling with weights proportional to the implied predictive probability distribution and, secondly, propagating the next latent mixture allocation and implicitly sampling the next particle vector. We introduce the methodology in the context of finite mixture models, before extending to any nonparametric mixture model with an available predictive probability function and focusing, in particular, on Dirichlet Process mixture models. In addition, we show that the algorithm provides a natural estimate for sequential Bayes factors and can facilitate selection between competing dynamic models. The framework is illustrated with numerous real and simulated data examples. (This is joint work with Carlos Carvalho, Nicholas Polson and Matt Taddy)

Oct. 28, 2009

A Poisson-compound Gamma model for species richness estimation

Prof. Jiping Wang : 3 p.m. in SEO 636
Abstract Suppose D distinct species are observed from an infinite population consisting of N (unknown) distinct species. The estimate of N from popular nonparametric methods can be substantially biased downward, while parametric approaches assuming a smooth abundance curve in general lack robustness. In this paper we propose a Poisson-compound Gamma approach, where the species abundance distribution Q is modeled as a Gamma mixture. We first show a nesting property of the Gamma mixture model, under which an arbitrary finite Gamma mixture can be uniquely re-written as a new Gamma mixture with components sharing a unified shape parameter, and mixed in the mean parameter. Thereby Q can be estimated using nonparametric maximum likelihood method for any given a. We further propose a least-squares cross-validation procedure for choice of the shape parameter to attain the desired smoothness of Q while controlling the goodness of fit of the model. The competitive performance of the resulting N-estimator is demonstrated using numerical studies and newly arising genomic data.

Nov. 4, 2009

Nested Latin Hypercube Designs

Prof. Peter Qian : 3 p.m. in SEO 636
Abstract We introduce a new type of design, called nested Latin hypercube design, for sequential integration and multi-fidelity computer modeling. A nested Latin hypercube design is defined to be a special Latin hypercube design that contains a smaller Latin hypercube design as a subset. Such designs are constructed by exploiting nested structures in random permutations. The constructed designs are also useful for solving stochastic optimization problems, including stochastic programs, the Monte Carlo EM algorithm and chance-constraint problems.

Nov. 11, 2009

Asymptotics of Maximum Partial Likelihood Estimators in General Semiparametric Multiplicative Hazard Models Under First Order Differentiability

Prof. Hanxiang Peng : 3 p.m. in SEO 636
Abstract In this talk, we discuss the asymptotic properties of a semiparametric multiplicative hazard model when the relative risk is expressed as a first order continuously differentiable parametric function. We show that the log- the partial likelihood function of the model is locally concave for an arbitrary continuously differentiable relative risk under suitable conditions. Then we derive the existence and uniqueness of the MPLE and show consistency. Using the convexity lemma and characterization of minimizers, we demonstrate that the MPLE of the parameter is asymptotically normal. As an application, we exhibit that the MPLE of the parameter in a model in which the log- the relative risk is expressed as a free-knot spline with knots in covariates uniquely exists in a neighborhood of the true parameter value and is consistent and asymptotically normal. In particular, we derive the asymptotic normality of the MPLE of the parameter in a model in which the log- relative risk is expressed as a free-knot quadratic spline which has first order continuous derivative.

Nov. 18, 2009

Random-effect Poisson Regression Analysis of Adverse Event Reports: The Relationship Between Antidepressants and Suicide

Prof. Dulal Bhaumik : 3 p.m. in SEO 636
Abstract A new statistical methodology is developed for analysis of spontaneous adverse event reports from post-marketing drug surveillance data. The method involves both empirical Bayes and fully-Bayes estimation of rate multipliers for each drug within a class of drugs, for a particular adverse event, based on a mixed-effects Poisson regression model. Both parametric and semi-parametric models for the random effect distribution are examined. The method is applied to data from FDA`s Adverse Event Reporting System (AERS) on the relationship between antidepressants and suicide. We obtain point estimates and 95% confidence intervals for the rate multiplier for each drug (e.g., antidepressants), which can be used to determine if a particular drug has an increased risk of association with a particular adverse event (e.g., suicide). Confidence intervals that do not include 1.0 provide evidence for either significant protective or harmful associations of the drug and the adverse effect. We also examine empirical Bayes, parametric Bayes and semi-parametric Bayes estimators of the rate multipliers and associated confidence intervals. Results of our analysis of the FDA AERS data revealed that newer antidepressants are associated with lower rates of suicide. This finding contradicts previous findings of FDA that newer antidepressants are causally related to increased suicidal thinking in children and young adults. Finally, we suggest changes in the AERS system to improve our ability to discover these adverse events.

Jan. 15, 2010

Efficient Crossover Designs and Limit Theories for Sample Covariances of Long-Memory Linear Processes

Wei Zheng : 2 p.m. in SEO 612
Abstract In crossover designs, it suffices to consider a linear model with effects of subjects, periods, direct and first-order carryover effects of treatments in most applications. We allowed the subject effects to be random, and thus included the fixed subject effects model as a special case. Efficient designs are proposed under this more flexible model. We also studied statistical properties and methods of constructions of these designs. For the remaining few minutes, I will talk about asymptotic behavior of sample autocovariances of long-memory linear processes. Also, some of my future interests will be addressed at the end.

Jan. 27, 2010

D-OPTIMAL DESIGNS FOR COMPLEX NONLINEAR MODELS

Ying Zhou : 3 p.m. in SEO 636
Abstract D-optimal designs for nonlinear chemical kinetics model and 2n-compartment models are investigated. We develop a method to obtain lower bounds of the maximal numbers of D-optimal design points to facility the computational search for the optimal design points in practice. We also investigate the conditions when the D-optimal design is the saturated design for bounded and unbounded design spaces. For each model discussed, the D-efficiency when the parameter misspecification happens is discussed.

Feb. 3, 2010

Optimal Allocation of Subjects to Multiple Treatments with Weibull Models

Cuilan Zhang : 3 p.m. in SEO 636
Abstract In clinical trials, it is important to find optimal allocation strategies to control the trials and benefit the participants. In our paper, we obtained explicit optimal allocation strategies for multiple treatments(2 and 3 treatments), in order to minimize expected total number of responses larger than a threshold, while guaranteeing a minimum power for the hypothesis that all treatments follow the same distribution. The assumed distributions are the standard Weibull models with common shape parameter. This optimal allocation strategy can be implemented by Doubly- biased coin design sequentially.

Feb. 17, 2010

Almost Sure Limit of the Smallest Eigenvalue of Some Sample Correlation Matrices

Han Xiao : 3 p.m. in SEO 636
Abstract Let $X^{(n)}=(X_{ij})$ be a $p \times n$ data matrix, where the $n$ columns form a random sample of size $n$ from a certain $p$-dimensional distribution. Let $R^{(n)}=(\rho_{ij})$ be the $p \times p$ sample correlation coefficient matrix of $X^{(n)}$; and $S^{(n)} = (1/n)X^{(n)}\left(X^{(n)}\right)^{\ast}-\bar{X}\bar{X}^{\ast}$ be the sample covariance matrix of $X^{(n)}$, where $\bar{X}$ is the mean vector of the $n$ observations. Assuming that $X_{ij}$'s are independent and identically distributed with finite fourth moment, we show that the smallest eigenvalue of $R^{(n)}$ converges almost surely to the limit $(1-\sqrt{c}\,)^2$ as $n \rightarrow \infty$ and $p/n \rightarrow c \in (0\,,\,\infty)$. We accomplish this by showing that the smallest eigenvalue of $S^{(n)}$ converges almost surely to $(1-\sqrt{c}\,)^2$.

Feb. 24, 2010

A Comparison Model for Measuring Individual Agreement via GEE Approach

Yuqing Tang : 3 p.m. in SEO 636
Abstract Evaluating individual agreement between different raters are often of interest in method comparison studies and in reliability studies. We propose a general comparison model and create a exible setting such that any subset of raters can be selected as test or reference raters and the individual agreement between them can be assessed. Two comparative agreement indices, the individual difference ratio (IDR) and within difference ratio (WDR), are proposed. IDR is a non-inferiority assessment such that the individual reading from different raters can not be inferior to the replicated readings within the same rater. When there is one test rater and one reference rater, our approach degenerates to FDA's method for evaluating individual bioequivalence under relative scale. WDR is a superiority assessment such that the precision of selected test raters can be better than that of selected reference raters. GEE approach is used for estimation and inference. Simulation study is conducted to assess the performance of our approach and the result shows that our method works well for both continuous data and categorical data.

March 3, 2010

Inference methods in functional mixed-effects models

Prof. David Degras : 3 p.m. in SEO 636
Abstract This work deals with the construction of simultaneous confidence bands (SCB) in functional mixed-effects models. SCB are useful graphical and analytic tools in data exploration, model construction, estimation of fixed and random effects, covariance estimation, prediction, and inference. The model under study can handle dummy variables and functional predictors, and its nonparametric covariance structure accounts for dependence in the data in a more flexible way than other models based on smoothing spline or wavelet representations. The SCB-based method allows to test local and nonparametric alternatives on the model components as opposed to the usual F-type tests. Some asymptotic theory is derived and a numerical study is presented to compare our method with other approaches currently in practice. We also illustrate our methodology by applying it to a real data set.

March 17, 2010

Model Selection of Correlation Structure for Clustered Data

Prof. Annie Qu : 3 p.m. in SEO 636
Abstract Model selection of correlation structure is a challenging problem because it involves a higher order of moments than model selection of covariates only. In addition, the high dimension of the correlation parameters could make the estimation of those parameters unreliable since the number of repeated measurements might be relatively small compared to the dimension of the correlation parameters. However, the correct specification of the correlation structure plays an important role in improving estimation efficiency for clustered data. We propose to select the correlation structure for clustered data from a number of candidate structures through a group-wise basis matrices selection strategy. The proposed method has the advantages of not requiring the likelihood function and of being computationally efficient. Also, the method can identify complex correlation structures. Furthermore, it is applicable for both continuous and discrete response data. In theory, we show that the proposed method enjoys the oracle property of selecting the true correlation structure consistently and estimating the correlation parameters with the same asymptotic normal distribution as if the true structure is known. This is joint work with Jianhui Zhou of University of Virginia.

March 31, 2010

The HYBRID GARCH Class of Models

Prof. Fangfang Wang : 3 p.m. in SEO 636
Abstract We propose a general GARCH framework that allows the use of different frequency returns to model conditional heteroskedasticity. We call the class of models High FrequencY Data-Based PRojectIon-Driven GARCH models as the GARCH dynamics are driven by what we call HYBRID processes. We study three broad classes of HYBRID processes: (1) parameter-free processes that are purely data-driven, (2) structural HYBRIDs where one assumes an underlying DGP for the high frequency data and finally (3) HYBRID filter processes. We develop the asymptotic theory of various estimators and study their properties in small samples via simulations. This is joint work with Eric Ghysels (University of North Carolina at Chapel Hill) and Xilong Chen (SAS Institute Inc.).

April 7, 2010

Regularized REML for Estimation and Selection of Fixed and Random Effects in Linear Mixed-Effects Models

Prof. Sijian Wang : 3 p.m. in SEO 636
Abstract The linear mixed effects model (LMM) is widely used in the analysis of clustered or longitudinal data. In the practice of LMM, inference on the structure of random effects component is of great importance not only to yield proper interpretation of subject-specific effects but also to draw valid statistical conclusions. This task of inference becomes significantly challenging when a large number of fixed effects and random effects are involved in the analysis. The difficulty of variable selection arises from the need of simultaneously regularizing both mean model and covariance structures, with possible parameter constraints between the two. In this paper, we propose a novel method of regularized restricted maximum likelihood to select fixed and random effects simultaneously in the LMM. The Cholesky decomposition is invoked to ensure the positive-definiteness of the selected covariance matrix of random effects, and selected random effects are invariant with respect to the ordering of predictors appearing in the model. We develop a new algorithm that solves the related optimization problem effectively, in which the computational load turns out to be comparable with that of the Newton-Raphson algorithm for MLE or REML in the LMM. We also investigate large sample properties for the proposed estimation, including the oracle property. Both simulation studies and data analysis are included for illustration. This is a joint work with Peter XK Song and Ji Zhu.

April 9, 2010

Social Network Models for Identifying Active Brain Regions from fMRI Data

Prof. Abhyuday Mandal : 2 p.m. in SEO 612
Abstract Functional magnetic resonance imaging (fMRI) is an important tool for scientists studying brain function. FMRI data are complex in nature: they are massive in size and a low signal-to-noise level makes the elimination of some noise prior to model fitting desirable for improved identification of true brain activity. We propose two methods of reducing this noise: generalized indicator functional analysis and a hidden Markov model. Brain regions showing increased fMRI signal while subjects engaged in a visual/spatial motor task are identified using concepts from social network analysis and statistical mechanics. Conditional probabilities of activation given the degree to which pairs of voxels are related are modeled for three groups: people with schizophrenia, their asymptomatic relatives, and control subjects. We compare the conditional probability maps obtained for each group to evaluate for between-group differences in extent of task-related signal. (Joint research with Ana M. Bargo, Lynne Seymour, Jennifer McDowelly, and Nicole A. Lazar)

April 14, 2010

ON THE ANALYSIS OF NON-NORMAL NONLINEAR TIME SERIES

Prof. Noelle Samia : 3 p.m. in SEO 636
Abstract The open-loop Threshold Model proposed by Tong (1990) is a stochastic piecewise-linear regression model useful for modeling conditionally normal response time-series data. How- ever, in many applications, the response variable is conditionally non-normal, e.g. Poisson or binomially distributed. We generalize the open-loop Threshold Model by introducing the Generalized Threshold Model (GTM). Specifically, it is assumed that the conditional probability distribution of the response variable belongs to the exponential family, and the conditional mean response is linked to some piecewise-linear stochastic regression function. The consistency and limiting distribution of the maximum likelihood estimator are derived. We illustrate the GTM with a real application on the annual number of human bubonic plague cases in Kazakhstan. This is based on joint work with Professor Kung-Sik Chan (University of Iowa) and Professor Nils C. Stenseth (University of Oslo).

April 21, 2010

Statistical issues in cancer-related copy number profiling

Prof. Hongmei Jiang : 3 p.m. in SEO 636
Abstract Copy number changes, either amplification or deletion of DNA materials, have been linked with cancer and other diseases. High-throughput DNA arrays enable simultaneous measurements of copy number on the genome-wide level. Given the intensity measurements from a single sample, methodologies and algorithms based on diverse techniques such as Hidden Markov Models, binary segmentation, and mixture models, have been developed to divide the genome into segments or regions of equal copy number. In this talk we will discuss how to find regions of recurrent gains or losses of DNA fragments across multiple cancer samples; how to identify chromosome regions which are associated with clinical variables such as relapse and survival outcome. The methods will be demonstrated on a liver cancer data set.

April 23, 2010

Optimal Design for Experiments with a Control

Prof. John Morgan : 4:15 p.m. in SEO 636
Abstract Standard optimality arguments for designed experiments rest on the assumption that all treatments are of equal interest. A notable exception is found in the ``test treatment versus control" (TvC) literature, where the control is allocated special status. Optimality work there has focused on all pairwise comparisons with the control, making no explicit account of how well test treatments are compared to one another. If the latter are also of consequence, it would be preferable to choose a design reflecting the relative importance placed on contrasts involving the control to that placed on contrasts of test treatments only. This talk develops the \emph{weighted} optimality approach for situations such as this, so that design selection may better reflect experimenter goals. When evaluating designs for comparing $v$ treatments, the basic idea is to assign weights $w_1,\ldots,w_v$ ($\sum_iw_i=1$) to account for differential treatment interest. In experiments with a control, and equal interest in the test treatments, this means weight $w_1$ is assigned to the control, and weight $w_2=(1-w_1)/(v-1)$ to each test treatment. These weights enter the evaluation through optimality measures, leading to, for example, weighted versions of the popular A, E, and MV measures of design efficacy. Families of weighted-optimal designs are identified under these criteria. Compared to their unweighted versions, it is shown that they less frequently agree on the best design. The classical approach in TvC design is shown to be a limiting case of the theory developed here.

April 28, 2010

Estimating the Accuracy of Verdicts in Criminal Trials when Truth Is Unknown

Prof. Bruce D. Spencer : 3 p.m. in SEO 636
Abstract Criminal trials may be viewed as complex classification procedures where the verdict represents classification as guilty or not guilty. Assessing the accuracy of verdicts is difficult because the "true" state of the defendant typically is unknown, and those cases where it is known are atypical. Yet, average accuracy of verdicts in criminal cases can be studied systematically and empirically provided we can obtain a second (or even a third) rating of the verdict. For example, in a jury trial the judge can also be asked for a verdict, as in the National Center for State Courts (NCSC) study of criminal cases from four jurisdictions in 2000-01. That study, like the famous Kalven-Zeisel study of the 1950s, showed only modest agreement between the judge and jury. Estimates of overall accuracy of verdicts are easily developed from the judge-jury agreement rate, and under plausible conditions the estimates of accuracy are optimistic. Estimates of false conviction rates and false acquittal rates, are more challenging, and are developed for the NCSC data with the use of log-linear latent class models. Those models, as well as models based on more than two raters, depend on stronger assumptions than the estimates of overall accuracy based on agreement rates. Numerical estimates of verdict accuracy are presented for the NCSC data and sources of uncertainty in the estimates are discussed, with particular attention to the effect of invalidity of the latent class. The estimates of the false conviction rates and false acquittal rates lead to questions about the appropriate balance of errors. Limitations of statistical decision theory for finding an optimal balance will be discussed.

Aug. 25, 2010

Organizational meeting

Junhui Wang : 3 p.m. in SEO 636

Sept. 8, 2010

Sparse Distance Weighted Discrimination and the Oracle Theory

Lingsong Zhang : 3 p.m. in SEO 636
Abstract Distance Weighted Discrimination (DWD) has recently been proposed as an attractive classification method. In this paper, we first show Fisher consistency of the DWD method, which justifies its use when there are sufficient data. However, the DWD classifier is not sparse, which makes the interpretation and prediction performance less attractive. We propose several sparse DWD methods, which incorporate variable selection techniques in classification using penalized loss functions to estimate the true hyperplane. We show that when an appropriate penalty is used, the sparse DWD method is consistent and the estimated normal vector has the oracle property under suitable conditions. We evaluate the finite sample performance of the proposed methods using simulations and illustrate the methods with an application to the Faroe island proteomic biomarker data.

Sept. 15, 2010

A Tale of Two Manifolds

Sayan Mukherjee : 3 p.m. in SEO 636
Abstract The focus is on the problem of supervised dimension reduction (SDR). We first formulate the problem with respect to the inference of a geometric property of the data, the gradient of the regression function with respect to the manifold that supports the marginal distribution. We provide an estimation algorithm, prove consistency, and explain why the gradient is salient for dimension reduction. We then reformulate SDR in a probabilistic framework and propose a Bayesian model, a mixture of inverse regressions. In this modeling framework the Grassman manifold plays a prominent role.

Sept. 22, 2010

Estimating rates at which books are mis-shelved

Hongmei Liu : 3 p.m. in SEO 636
Abstract Basic ideas of survey sampling were utilized in a project for STAT431 "Introduction to Survey Sampling" at the University of Illinois at Chicago, taught and supervised by Professor Hedayat in the fall 2009, to estimate the mis-shelving rate of books at the University of Illinois-Chicago (UIC) Daley library. We use this project as an example to illustrate how to design a sampling plan, sample size, and how to estimate rates at which books are mis-shelved.

Sept. 29, 2010

Curvature, Robustness and Optimal Design in Applied Generalized Nonlinear Regression Modelling

Timothy E. O'Brien : 3 p.m. in SEO 636
Abstract Researchers often find that nonlinear regression models are more applicable for modelling various biological, physical and chemical processes than are linear ones since they tend to fit the data well and since these models (and model parameters) are more scientifically meaningful. These researchers are thus often in a position of requiring optimal or near-optimal designs for a given nonlinear model. A common shortcoming of most optimal designs for nonlinear models used in practical settings, however, is that these designs typically focus only on (first-order) parameter variance or predicted variance, and thus ignore the inherent nonlinear of the assumed model function. Another shortcoming of optimal designs is that they often have only p support points, where p is the number of model parameters. Measures of marginal curvature, first introduced in Clarke (1987) and further developed in Haines et al (2004), provide a useful means of assessing this nonlinearity. Other relevant developments are the second-order volume design criterion introduced in Hamilton and Watts (1985) and extended in O'Brien (1992, 2010), and the second-order MSE criterion developed and illustrated in Clarke and Haines (1995). This talk examines various robust design criteria and those based on second-order (curvature) considerations. These techniques, coded in the GAUSS and SAS/IML software packages, are illustrated with several examples including one from a preclinical dose-response setting encountered in a recent consulting session.

Oct. 6, 2010

Dynamic Cluster-based Sliced Inverse Regression for Forecasting Microeconomics Variables

Yue Yu : 3 p.m. in SEO 636
Abstract Sliced Inverse Regression (SIR) proposed by Ker-Chau Li (1991) is a widely used semiparametric technique to reduce the dimensions of regression problems. But the microeconomics data are time dependent and usually highly correlated. In our study, we use clustering methods along with SIR, in order to reduce the multicollinearity and the difficulty of choosing the dimensions. And the dynamic version of cluster SIR methods is used to analyze the autoregressive model of the microeconomics data. Our simulation result shows that the dynamic cluster SIR is superior, comparing with the empirical accuracy of all the models in Stock and Watson's paper (2005) for forecasting U.S. macroeconomic time series over a 30-year period.

Oct. 13, 2010

Joint Modeling of Longitudinal and Time to Event Data with Random Changepoints

Yuan Xu : 3 p.m. in SEO 636
Abstract We develop a joint model of longitudinally observed cognitive data and survival data to the onset of dementia. We incorporate latent random change points in the model representing an accelerated cognitive decline prior to the onset of dementia. We aim to investigate how different covariates of subjects, such as baseline age, education and genetic risk factors, affect the timing of cognitive decline acceleration. We also assess how different groups of subjects behave on cognitive decline before and after the change point. The model combines a longitudinal mixed effects model with a Cox proportional hazards model connected by a random change point with a log normal distribution. The parameters are estimated by the maximum likelihood method through an ECM algorithm. Compared with joint models with change points developed previously by other authors, our model has several advantages. First, our model uses the semi-parametric Cox model instead of a parametric model for the survival data, therefore is more flexible to different survival distributions. Second, we use the maximum likelihood method and an ECM algorithm to estimate the parameters to avoid the prior assumptions on model parameters. Third, we propose a compromised Fisher information method other than profile likelihood method to obtain a better estimation of the standard errors of the MLEs for model parameters. Finally, the proposed model is successfully implemented to study the preclinical acceleration on the rate of cognitive decline as well as its implication on the risk of developing dementia.

Oct. 20, 2010

Prediction of Travel Times for CTA Buses

Troy Hernandez : 3 p.m. in SEO 636
Abstract Bus arrival prediction times are useful to many passengers of public transportation. There are many features of the environment that can be used to predict bus arrival times and many possible representations of these features. Two representations used to predict bus arrival times are discussed: a memoryless representation that uses only the current state of the bus and a full-memory or trajectory-based representation that uses the full history of the bus run.

Oct. 27, 2010

High dimensional classification and its application in pharmacogenomics research

Jun Xie : 3 p.m. in SEO 636
Abstract Many statistical classification methods, e.g., Fisher's linear discriminant analysis, cannot be directly applied to high dimensional data, where the number of variables is larger than the sample size. While high dimensional data analysis has been broadly discussed in statistics community, the impact of dimensionality on classifications is poorly understood. We examine and compare high dimensional classification methods through an application in pharmacogenomics research, where high-dimensional gene expression microarray data are used to predict patients' responses to a drug. Compared with most gene expression classification studies to detect strong signals, for instance tumor versus normal, a classifier between patients' response and non-response is more challenging and may be nonlinear. We introduce several new classification methods, including a sparse linear discriminant method, random projection, and a distribution based classification involving second-order interactions, as potential tools to deal with high dimensionality. We also want to call attentions to theories of high dimensional classification, where there are only few results available.

Nov. 3, 2010

On Multiple Testing and the Monotone Likelihood Ratio Condition

Hongyuan Cao : 3 p.m. in SEO 636
Abstract High-throughput screening has become an important mainstay for con- temporary biomedical research. A standard approach is to get p-values and adjust for multiple comparison in a manner that controls false discovery rate (FDR). The concavity of p-value distribution under the alternative has been a standard condition for developing many FDR procedures: Storey (2003), Genovese and Wasserman (2004), Kosorok and Ma (2007). A more general concept is the monotone likelihood ratio condition (MLRC) introduced in Sun and Cai (2007). We show in this paper that the concavity assumption can be violated for (i) a simple heteroscedastic normal mixture model and (ii) dependent tests. Some interesting implications, including different testing procedures (step-up vs step-down), the choice of test statistic and the power definition in multiple testing are discussed. This is joint work with Wenguang Sun and Michael R. Kosorok.

Nov. 10, 2010

Maximum Likelihood Degree for the Random Effects Mode

Elizabeth Gross : 3 p.m. in SEO 636
Abstract Maximum likelihood estimation is a common problem explored in Algebraic Statistics, a field that focuses on the applications of algebraic geometry to the study of statistical models. In maximum likelihood estimation if the likelihood equations are algebraic, the maximum likelihood degree (ML degree) is a measure of the algebraic complexity of the estimation problem. It is the degree of the variety characterized by the system of likelihood equations, or, equivalently, the number of complex solutions of the system for generic data. The ML degree is specific to the statistical model and there are only two other classes of models for which an explicit formula for the ML degree is known. In this talk, we will look at the analysis of variance model with random effects and give an explicit formula for the ML degree. We also explore the number of feasible (real, positive) solutions computationally. This is joint work with Mathias Drton and Sonja Petrovic.

Nov. 17, 2010

Efficient Crossover Designs under Subject Dropout

Dibyen Majumdar : 3 p.m. in SEO 636
Abstract Crossover studies are used in different areas of statistical applications and there is a substantial literature that focus on identifying and constructing efficient designs for these studies. However, even well-designed crossover studies often lose their statistical properties if subjects drop out before the end of the study. We will explore the problem of subject dropout and the effect on properties of the design, and search for efficient designs that are robust to subject dropout.

Nov. 23, 2010

Estimation and Variable Selection for Generalized Additive Partial Linear Models

Li Wang : 1:30 p.m. in SEO 636
Abstract We study a class of generalized additive partial linear models. We propose the use of polynomial spline smoothing for estimation of nonparametric functions, and derive the quasi-likelihood based estimators for the linear parameters. We establish asymptotic normality for the estimators of the parametric components. The procedure avoids solving big system of equations as in kernel-based procedures and thus results in gains in computational simplicity. We further develop a class of variable selection procedures for the linear parameters by employing a nonconcave penalized likelihood, which is shown to have an oracle property. Monte Carlo simulations and an analysis of a dataset from Pima Indian diabetes study are presented for illustration.

Dec. 1, 2010

Kernel density estimation for time series: an asymptotic theory

Yinxiao Huang : 3 p.m. in SEO 636
Abstract In this paper, we give a unified analysis of both the nonparametric kernel density estimator and regression estimator under our dependence structure. Asymptotic results such as uniform convergence rate, $L^p$ convergence rate, asymptotic normality are obtained under fairly mild conditions. In particular, we allow certain long memory processes as well. A closely related problem, the recursive kernel estimator where the bandwidth changes with each observation, is still in progress and I hope to get it done pretty soon. Therefore, I may talk about the recursive estimator as well if time permits.

Feb. 9, 2011

Cluster analysis of high-dimensional data via regularized K-means

Wei Sun : 3 p.m. in SEO 636
Abstract K-means clustering is a widely used tool for cluster analysis due to its conceptual simplicity and computational efficiency. However, its performance can be distorted when clustering high-dimensional data where the number of variables becomes relatively large and many of them may contain no information about the clustering structure. In this talk we will discuss a novel high-dimensional cluster analysis method via regularized k-means clustering, which can simultaneously cluster similar observations and eliminate redundant variables. The key idea is to formulate the k-means clustering in a form of regularization, with an adaptive group lasso penalty term on cluster centers. Then we will talk about the selection criterion based on clustering stability to optimally balance the trade-off between the clustering model fitting and sparsity. The effectiveness of the proposed method is demonstrated through a variety of numerical experiments as well as applications to two gene microarray examples.

Feb. 16, 2011

Optimal Designs for Two-Level Factorial Experiments with Binary Response

Abhyuday Mandal : 3 p.m. in SEO 636
Abstract We consider the problem of obtaining locally D-optimal designs for factorial experiments with qualitative factors at two levels each with binary response. For the 2^2 factorial experiment with main effects model we obtain optimal designs analytically in special cases and demonstrate how to obtain a solution in the general case using Cylindrical Algebraic Decomposition. We also study the sensitivity of the D-optimal designs to misspecification of the assumed parameter values. When there is no basis to make an informed choice of the assumed values, we recommend the use of the uniform design. For the general 2^k case we show that the uniform design has a maximin property.

Feb. 23, 2011

Classification Based on Permanent Process with Cyclic Approximation

Jie Yang : 3 p.m. in SEO 636
Abstract We propose a stochastic classification model based on a permanent process. Unlike many research works in the literature, the proposed model assumes only exchangeability instead of independence on observations. Regardless of the number of classes or the dimension of the feature variables, the model may require only 2-3 parameters for fitting the covariance structure within clusters. It works well even if the class occupies non-convex, disjoint regions, or regions overlapped with other classes in the feature space. The proposed model requires calculation of ratios of weighted permanents, which is an NP-hard problem. We propose a series of approximations for weighted permanent ratio based on cyclic expansions. The classification based on cyclic approximations works reasonably well.

March 2, 2011

How a drug is made-R&D functions in pharmaceutical companies

William Zhao : 3 p.m. in SEO 636
Abstract An Overview of Pharmaceutical R&D and Roles of Statisticians. The drug research and development process will be introduced. Examples will be provided. Experiences will be shared on the roles of statisticians in the process.

March 16, 2011

High Order Laplace Approximation for Design and Statistical Inference of Dependent Data

Zhengyuan Zhu : 3 p.m. in SEO 636
Abstract Laplace approximation has been used in many statistical applications to approximate integrals. In this talk we present several applications which use high order Laplace approximation to derive theoretical results and develop efficient algorithms, including construction of prediction interval which has zero second order coverage probability bias, design criteria for prediction with estimated covariance parameters, theoretical comparison of predictive densities for dependent observations, and approximate inference for spatial generalized linear mixed models.

March 30, 2011

Statistical Issues in Drug Safety: The curious case of Antidepressants, Anticonvulsants, ...., and Suicide

Robert Gibbons : 3 p.m. in SEO 636
Abstract In 2003, the U.S. FDA, MHRA in the U.K., and European union released public health advisories for a possible causal link between antidepressant treatment and suicide in children and adolescents ages 18 and under. This led the U.S. FDA to issue a black box warning for antidepressant treatment of childhood depression in 2004, which was later extended to include young adults (18-24) in 2006. Following these warnings, rather than observing the anticipated decrease in youth suicide rates, record increases in youth suicide rates were observed in both the U.S. and Europe. In this presentation, we review the data and statistical methodology that led to the public health advisories and black box warning, and the data that led to the record increases in youth suicide rates and discuss their possible relationship. New statistical and experimental design approaches to post-marketing drug safety surveillance are developed, discussed and illustrated.

April 6, 2011

NSF funding opportunities in mathematical sciences

Haiyan Cai : 3 p.m. in SEO 636
Abstract In this talk I plan to give a brief introduction of the Division of Mathematical Sciences (DMS) of NSF and its funding opportunities, including its investment goals, investment areas, disciplinary and interdisciplinary programs, and proposal evaluation criteria. I will also be happy to answer questions.

April 13, 2011

Analysis of Large-Scale Computer Experiments

Lulu Kang : 3 p.m. in SEO 636
Abstract Computer experiments simulate the engineering systems by implementing the mathematical models governing the systems in computers. Recently, experiments having large number of input variables and experimental runs started to emerge. In the existing literature, kriging has been commonly used for approximating the complex computer models, but it has limitations for dealing with the large-scale experiments due to its computational complexity and numerical stability. In this work, I propose a new modeling approach known as regression-based inverse distance weighting (RIDW). The new predictor is shown to be computationally more efficient than kriging while producing comparable prediction performance. We also develop a heuristic method for constructing confidence intervals for prediction. I will also discuss extensions of RIDW and my future research directions on this exciting topic.

April 20, 2011

Optimal designs under variety constrains

Min Yang : 3 p.m. in SEO 636
Abstract In many practical studies, we need to consider the optimality problem under variety constrains. For example the cost of experiment may depend on the experiment point we choose and we need to control the total cost. On the other hand, we may want to make sure the design is at least efficient at certain level under different optimality criterion. In this talk, we will discuss such question.

April 27, 2011

Godambe's Non-existence Theorem - Revisited & Refined Through Matrix Arguments

Bikas K. Sinha : 3 p.m. in SEO 636
Abstract In the context of a finite labeled population, for estimation of a population total or mean, there does not exist any umvue, unless the sampling design is trivially a cluster sampling design. This is the celebrated Godambe's non-existence theorem. Basu established another version of it. We revisit the proof and strengthen the result by using matrix arguments.

May 4, 2011

Analysis of Covariance under Inverse Gaussian Model

Reza Meshkani : 3 p.m. in SEO 636
Abstract In this talk, the limitations of normal model for Analysis of Covariance for positive right-skewed variables are considered. Specifically, an Inverse Gaussian variable is considered whose variance depends on its mean thus violating the usual assumptions of Normal linear model. Instead of appealing to transformations which makes interpretations of the results awkward, we propose a method of direct statistical analysis from both Maximum Likelihood and Bayesian perspectives. The formulas for adjusting treatment effects are given and their properties are discussed. To provide explicit formulas, conjugate priors are considered. The posterior distributions are derived and procedures of adjustment for covariates are presented.

May 5, 2011

Eigenproblem and permanents over max algebra

Ravindra Bapat : 3 p.m. in SEO 1227
Abstract The max algebra consists of the set of real numbers, along with negative infinity, equipped with two binary operations, maximization and addition. The algebra is useful in describing certain conventionally nonlinear systems in a linear fashion. We discuss basic aspects of the eigenproblem in max algebra, pointing out that it can be seen as a limiting case of the Perron-Frobenius theory for nonnegative matrices. We also discuss properties of the permanent over the max algebra, which is the same as the maximum value of the classical assignment problem.

Aug. 3, 2011

Understanding Species' Abundance

Prof Bikas K Sinha : 4 p.m. in SEO 636
Abstract Considered is a multi-species assemblage in an infinite population with unknown and possibly heterogeneous species' abundance levels. A random sample of a fixed size (n) has been drawn only to realize a certain species distribution. At this stage, there are two interesting inference problems : (i) Prediction of the unknown proportion of the collective abundance of all hitherto unrealized species; (ii) Assessment of the quantum of additional units to be sampled [in terms of the sample size n] to realize a certain number of hitherto unobserved species. Finite population analogue of this problem is also worth discussing. I propose to review the available literature in this fascinating area of research.

Aug. 24, 2011

Statistics Seminar Organization Meeting

Statistics Group : 4 p.m. in SEO 636

Aug. 31, 2011

Search Designs for Model Selections in Fractional Factorial Experiments

Kashinath Chatterjee : 4 p.m. in SEO 636
Abstract The main objective of this presentation is to introduce the notion of Search Design, pioneered by Srivatava (1975), and its application to model selection in fractional factorial experiments.

Sept. 7, 2011

A Bayesian test of normality versus a Dirichlet process mixture alternative

Ryan Martin : 4 p.m. in SEO 636
Abstract Testing if a p-dimensional sample, for p >= 1, comes from a normal population is a fundamental problem in statistics. In this talk I will describe a new Bayesian test of p-variate normality against an alternative hypothesis characterized by a certain Dirichlet process mixture model. I will show that this nonparametric alternative satisfies the desirable embedding and predictive matching properties with respect to the normal null model. To compute the Bayes factor, an efficient sequential importance sampler is is proposed for evaluating the marginal likelihood under the nonparametric alternative. Numerical examples demonstrate that the proposed test has satisfactory discriminatory power when the distribution is not normal, and does not tend to over-fit when the distribution is normal.

Sept. 14, 2011

Concentration property and Log-Sobolev inequality for SDE's driven by fractional Brownian motions

Cheng Ouyang : 4 p.m. in SEO 636
Abstract Stochastic differential equations (SDE) driven by various random processes are important subject in both probability theory and applications, as they provide mathematical models for systems that evolve under random forces. Among them, study of SDE's driven by fractional Brownian motions is an active area in current research. In the talk, I will first give a brief introduction to this topic, and then present two resent results - namely, the concentration property and Log-Sobolev inequality - on the law of solutions to SDE's driven by fractional Brownian motions.

Sept. 21, 2011

Latent Class Analysis and Regression Analysis in Health Care Market Research / Mulitcategory Probability Estimation via Quantile Kernel Regression

Zhifan Zhang and Tu Xu : 4 p.m. in SEO 636
Abstract (By Zhifan Zhang) Latent Class Analysis (LCA) is used to find groups or subtypes of cases in multivariate categorical data. In Health care Market Research, Latent Class Analysis is used in consumer segmentation to identify discriminating variables that would enable the development of differentiated segments of customers to inform post-Reform Strategies. In this talk, we will discuss the usage of Latent Class Analysis and Regression analysis in Health care Market Research. (By Tu Xu) Multiclass probability estimation is an very important problem in statistics and data mining. The traditional probability estimation problem is commonly dealt by regression techniques such as multiple logistic regression, or the density estimation approaches such as linear discriminant analysis(LDA) and quadratic discriminant analysis. In this talk,we propose a new model-free method for estimating multiclass probabilities based quantile kernel regrssion will be introduced. It does not impose any strong parametric assumption on the underlying distribution and can be applied for a wide range of large-margin classification methods. Our simulations show great performance of our method compared with existing methods.

Sept. 28, 2011

An approach to modeling asymmetric multivariate spatial covariance structures

Bo Li : 4 p.m. in SEO 636
Abstract We propose a framework in light of the delay effect to model the asymmetry of multivariate covariance functions that is often exhibited in real data. This general approach can endow any valid symmetric multivariate covariance function with the ability of modeling asymmetry and is very easy to implement. Our simulations and real data examples show that asymmetric multivariate covariance functions based on our approach can achieve remarkable improvements in prediction over symmetric models.

Oct. 5, 2011

X-chromosome Genetic Association Analysis with Related Individuals

Mary Sara McPeek : 4 p.m. in SEO 636
Abstract Common diseases such as asthma, diabetes, and hypertension, which currently account for a large portion of the health care burden, are complex in the sense that they are influenced by many factors, both environmental and genetic. One fundamental problem of interest is to understand what the genetic risk factors are that predispose some people to get a particular complex disease. Despite the potential for complex traits to have X-linked causal genes, genetic association methods have primarily been developed for the analysis of markers on the autosomal chromosomes, and significantly less attention has been given to analyzing X-linked markers. We develop methods for case-control association testing of X-chromosome variants in samples in which some individuals are related. Our methods are applicable to association studies with completely general combinations of family and case-control designs, including large complex pedigrees. Even in the context of large complex pedigrees, the methods are computationally feasible for analysis of millions of variants. We allow for sex-specific prevalence, and we allow both unaffected controls and controls of unknown phenotype in the analysis. We discuss some of the distinct challenges posed by X-chromosome association analysis in contrast to autosomal association analysis. We discuss the performance of the methods in the context of several data sets as well as in simulations.

Oct. 12, 2011

Multivariate estimates for unsigned Gaussian functions

Ang Wei : 4 p.m. in SEO 636
Abstract Unsigned Gaussian functions (functions involving absolute value) arise in a variety of contexts, such as random polynomials and matrices. We apply integral representations and matrix analysis to estimate the absolute moments and quadratic forms of Gaussian vectors. We also present applications of these results to game theory, communication theory and astrophysics.

Oct. 19, 2011

Latent Space Model for Aggregated Relational Data

Tian Zheng : 4 p.m. in SEO 636
Abstract Aggregated Relational Data (ARD) are indirect network data collected using survey questions of the form "how many X's do you know?" It is most often used to estimate the size of populations that are difficult to count directly and allows researchers to choose specific subpopulations of interest without sampling or surveying members of these subpopulations directly. What has been under-utilized is the indirect information on social structure captured by ARD. In this talk, I present a latent space model and Bayesian computation framework for inference and estimation of social structures using ARD from non-network samples in social networks, the variation of social structures in subnetworks, and the relations between (hard-to-reach) subpopulations.

Oct. 25, 2011

Some aspects of stochastic differential equations driven by fractional Brownian motions

Fabrice Baudoin : 4 p.m. in SEO 636
Abstract In this talk we will review several results on stochastic differential equations driven by fractional Brownian motions that the speakers obtained in a series of more or less recent works. We shall in particular focus on the study of gradients bounds, Gaussian heat kernels bounds and small time asymptotics for the operators naturally associated with such equations. The presentation will be based on joint works with L. Coutin, M. Hairer, C. Ouyang and S. Tindel.

Oct. 26, 2011

Quantile tomography: using quantiles with directional data

Ivan Mizera : 4 p.m. in SEO 636
Abstract Directional quantile envelopes---essentially, depth contours---are a possible way to condense the directional quantile information, the information carried by the quantiles of projections. In typical circumstances, they allow for relatively faithful and straightforward retrieval of the directional quantiles, offering a straightforward probabilistic interpretation in terms of the tangent mass at smooth boundary points. They can be viewed as a natural, nonparametric extension of ``multivariate quantiles'' yielded by fitted multivariate normal distribution, and, as illustrated on data examples, their construction can be adapted to elaborate frameworks---like estimation of extreme quantiles, and directional quantile regression---that require more sophisticated estimation methods than simply evaluating quantiles for empirical distributions. Their estimates are affine equivariant whenever the estimators of directional quantiles are translation and scale equivariant; mathematically, they express the dual aspect of directional quantiles.

Nov. 2, 2011

Simultaneous Linear Quantile Regression: A Semiparametric Bayesian Approach

Surya Tapas Tokdar : 4 p.m. in SEO 636
Abstract I'll introduce a semi-parametric Bayesian framework for a simultaneous analysis of linear quantile regression models. A simultaneous analysis is essential to attain the true potential of the quantile regression framework, but is computationally challenging due to the associated monotonicity constraint on the quantile curves. For a univariate covariate, we present a simpler equivalent characterization of the monotonicity constraint through an interpolation of two monotone curves. The resulting formulation leads to a tractable likelihood function and is embedded within a Bayesian framework where the two monotone curves are modeled via logistic transformations of a smooth Gaussian process. A multivariate extension is proposed by combining the full support univariate model with a linear projection of the predictors. The resulting single-index model remains easy to fit and provides substantial and measurable improvement over the first order linear heteroscedastic model. I'll provide two illustrative applications to tropical cyclone intensity and birth weight.

Nov. 9, 2011

Transportation Cost Inequality for Gaussian Measure and Coupling of Brownian Motion

Elton Hsu : 4 p.m. in SEO 636
Abstract We will show how to use synchronizing Brownian motion to prove Talagrand's transportation cost inequality for the standard Gaussian measure. This proof can be readily generalized to a generalization of this inequality to the heat kernel measure on a Riemannian manifold using the synchronizing coupling of Riemannian Brownian motion.

Nov. 16, 2011

On Maximum Empirical Likelihood Estimation And Related Topics

Hanxiang Peng : 4 p.m. in SEO 636
Abstract In this talk, I will focus on maximum empirical likelihood estimation in the case of constraint functions that may be discontinuous and/or depend on additional parameters. The later is the case in applications to semiparametric models where the constraint functions may depend on the nuisance parameter. Our results are thus formulated for empirical likelihoods based on estimated constraint functions that may also be irregular. Applications of our results are discussed to inference problems about quantiles under possibly additional information on the underlying distribution and to partial adaption.

Nov. 30, 2011

An Interactive Allocation Rule for Networks

Surajit Borkotokey : 4 p.m. in SEO 636
Abstract We propose an allocation rule that takes into account the importance of players and their links. Since a network describes the interaction structure between agents, our allocation rule covers both bilateral and multilateral interactions. We provide a characterization of this rule in terms of well known axioms and compare it to other allocation rules in the literature.

Jan. 11, 2012

Completion of missing entries in matrices and tensors

Shmuel Friedland : 4:15 p.m. in SEO 636
Abstract In many instances in measuring multidimensional data, as matrices and tensors, one confronts the following problems: noisy data, missing entries and data reduction. There are many statistical and mathematical methods to deal with these problems. In this talk we survey some of the known methods and expand on the methods that the speaker was working on. A variant of this talk is available at http://homepages.math.uic.edu/$\sim$friedlan/complmattenSep11.pdf

Jan. 25, 2012

Yield Components and Growth Components Analysis for Soybeans // An Assessment of Mathematics Placement Policy Changes at UIC

Yan Sun, Ella Revzin : 4 p.m. in SEO 636
Abstract Yan Sun's Abstract: Soybean is an important crop in the US and around the world. So the analysis of soybean yield components as well as the growth components is always a hot topic in research. Our work originates from the requirement of a client and is based on the experiment data provided by him. The experiment has a split-split plot structure. Our task is to find significant factors for seed yield and protein concentration, which are both yield components. Also, we establish regression functions for some growth components (leaf area and leaf dry biomass) with respect to the growing time. Ella Revzin's Abstract: In the Fall 2011, the University of Illinois at Chicago changed its mathematics placement policy. Instead of placing students based on ACT scores and results on written tests, the mathematics department switched to using an online testing and learning software, ALEKS. This talk will summarize the results from analyzing student outcomes after and before the policy change; as well as an assessment of how effective the ALEKS system was in placing students in selected undergraduate courses.

Feb. 1, 2012

Jump robust two time scale covariance estimation and realized volatility budgets

Jin Zhang : 4:15 p.m. in SEO 636
Abstract We estimate the daily integrated variance and covariance of stock returns using high-frequency data in the presence of jumps, market microstructure noise and non-synchronous trading. For this we propose jump robust two time scale (co)variance estimators and verify their reduced bias and mean square error in simulation studies. We use these estimators to construct the ex-post portfolio realized volatility (RV) budget, determining each portfolio component's contribution to the RV of the portfolio return. These RV budgets provide insight into the risk concentration of a portfolio. Furthermore, the RV budgets can be directly used in a portfolio strategy, called the equal-risk-contribution allocation strategy. This yields both a higher average return and lower standard deviation out-of-sample than the equal-weight portfolio for the stocks in the Dow Jones Industrial Average over the period October 2007-May 2009.

Feb. 8, 2012

Assessment of Agreement in Linear/Generalized Linear Mixed Models

Yue Yu : 4 p.m. in SEO 636
Abstract Study of measuring agreement is mainly aimed to answer one question, whether the readings from one instrument/method agree with the ones from another instrument/method. In this talk, we are going to present a general method to assess agreement for a wide range of data types with repeated measurements using linear and generalized linear mixed models. Likelihood-based approaches are developed to estimate all the within- and between-instrument agreement statistics. and asymptotic properties of these agreement estimates are discussed for different data structures. Furthermore, our method has the merit of handling missing values and covariates naturally. And a new set of restricted agreement statistics is proposed in order to capture the true random variations and between-instrument effects rather than the covariate effects. Simulations and several case studies, involving method comparison and bioequivalence, are used to show the accuracy and effectiveness of our method.

Feb. 15, 2012

Graphical Representation of Biological Sequences and Its Applications

Chenglong Yu : 4 p.m. in SEO 636
Abstract Among all existing alignment-free methods for comparing biological sequences, the sequence graphical representation provides a simple approach to view, sort, and compare gene structures. The aim of graphical representation is to display DNA or protein sequences graphically so that we can easily find out visually how similar or how different they are. Of course, only the visual comparison of sequences is not enough for the follow-up research work. We need more accurate comparison. This leads us to develop the application of the graphical representation for biological sequences. I will talk about two contributions for this direction. (1) We construct a protein map with the help of our proposed new graphical representation for protein sequences. Each protein sequence can be represented as a point in this map, and cluster analysis of proteins can be performed for comparison between the points. This protein map can be used to mathematically specify the similarity of two proteins and predict properties of an unknown protein based on its amino acid sequence. (2) We construct a novel genome space with biological geometry, which is a subspace in R^N. In this space each point corresponds to a genome. The natural distance between two points in the genome space reflects the biological distance between these two genomes. The genome space will provide a new powerful tool for analyzing the classification of genomes and their phylogenetic relationships.

Feb. 21, 2012

Some Finer Aspects of de la Garza Phenomenon

Bikas Sinha : 4 p.m. in SEO 636
Abstract de la Garza Phenomenon relates to the Information Matrix in the context of a standard Gauss-Markov Linear Model. It works well in the framework of approximate or continuous designs. For discrete designs, one has to be careful in extracting its full spirit. We propose to discuss some features of this highly fascinating area of research.

Feb. 22, 2012

Penalized Linear Discriminant Analysis for Family Studies

Yixin Fang : 4 p.m. in SEO 636
Abstract In family studies with multiple continuous phenotypes, we are interested in finding linear combinations of the phenotypes with large heritabilities, which can be considered as new phenotypes for genetic analysis. The problem can be recast as linear discriminant analysis (LDA). When the number of phenotypes is large, LDA is not appropriate for two reasons: the standard estimate for the within-family covariance matrix is singular, and it is difficult to interpret the newly defined phenotypes. Here we propose a novel version of penalized LDA, with an $L_1^2$ penalty in the denominator of the Rayleigh quotient. Besides overcoming the above two problems, the proposed method has at least three advantages compared with the existing regularization methods. First, it solves the singularity problem and achieves the sparsity property simultaneously. Second, the method is scale-invariant. Third, the consistency can be proved. We evaluate the performances of the method using simulations and two family studies.

Feb. 29, 2012

Persistence of a Swamp Rabbit Metapopulation: The Incidence Function Model Approach / A Quantile Regression Study of Climate Change in Chicago

Jennifer Pajda-Delao / Julien Leider : 4 p.m. in SEO 636
Abstract Jennifer's Abstract: We evaluate the status and distribution of swamp rabbits (Sylvilagus aquaticus) in Missouri using the Incidence Function Model and logistic regression in an effort to assess the long term viability of the Missouri metapopulation. We used results of latrine surveys performed in 1992 and 2001 to estimate the likelihood of persistence of swamp rabbits over periods of 9 to 1000 years. Under current conditions, more than 50% of the patches are predicted to contain rabbits after 1000 years. Logistic regression revealed that both patch area and patch isolation were significantly related to patch occupancy, and play key roles in the incidence of swamp rabbits. Julien's Abstract: This study uses quantile regression combined with time series methods to analyze change in temperatures in Chicago during the period 1960-2010. It builds on previous work in applying quantile regression methods to climate data by Timofeev and Sterin (2010) and work by the Chicago Climate Task Force on analyzing climate change in Chicago. We use data from the Chicago O'Hare Airport weather station archived by the National Climatic Data Center to look at changes in weekly average temperatures. We use the method described by Xiao et al. (2003) to remove autocorrelation in the data, the rank-score method with IID assumption to calculate confidence intervals, and nonparametric local linear quantile regression to estimate temperature trends. We find that the decade 1960-1969 was significantly cooler than later decades around the middle of the yearly seasonal cycle at both the median and 95th percentile of the temperature distribution. However, we do not find a significant change across later decades of the study period.

March 7, 2012

On Optimal Harvesting Problems in Random Environments

Chao Zhu : 4 p.m. in SEO 636
Abstract We consider the optimal harvesting strategy for a single species living in random environments whose growth is given by a regime-switching diffusion. Harvesting acts as a (stochastic) control on the size of the population. The objective is to find a harvesting strategy which maximizes the expected total discounted income from harvesting up to the time of extinction of the species; the income rate is allowed to be state- and environment-dependent. This is a singular stochastic control problem with both the extinction time and the optimal harvesting policy depending on the initial condition. One aspect of receiving payments up to the random time of extinction is that small changes in the initial population size may significantly alter the extinction time when using the same harvesting policy. Consequently, one no longer obtains continuity of the value function using standard arguments for either regular or singular control problems having a fixed time horizon. We introduce a new sufficient condition under which the continuity of the value function for the regime-switching model is established. Further, it is shown that the value function is the unique viscosity solution of a coupled system of quasi-variational inequalities. We also establishes a verification theorem and, based on this theorem, an $\varepsilon$-optimal harvesting strategy is constructed under certain conditions on the model. This is a joint work with Qingshuo Song and Richard Stockbridge.

March 14, 2012

Compactly Supported Multivariate Covariance Modeling with Application to Spatial Taper

Juan Du : 4 p.m. in SEO 636
Abstract We derive several classes of covariance matrix functions whose entries are compactly supported. These compactly supported matrix functions are used as building blocks to formulate other covariance matrix functions for modeling of multivariate spatial processes. In particular, a multivariate version of the celebrated spherical model is produced, as well as a class of second-order multivariate stochastic processes whose direct and cross covariance functions are of Pólya type. On the other hand, by employing some of the proposed compactly supported correlation matrix functions as the tapering matrix function, we study the multivariate generalization of the spatial covariance tapering technique, which is useful to mitigate the numerical burdens in dealing with the large spatial data sets by making covariance matrices sparse. Simulation study is conducted to show the computational efficiency and application in spatial prediction by using proposed multivariate tapering technique.

March 28, 2012

Inverting Analytic Characteristic Functions

Liming Feng : 4 p.m. in SEO 636
Abstract Analytic characteristic functions naturally arise in financial engineering applications. We explore the analyticity of such characteristic functions and propose simple but highly accurate inversion schemes. The schemes have the following advantages: (1) they are very easy to implement; one does not need to rely on commercial numerical packages; (2) despite the simplicity, they are highly accurate, with exponentially decaying errors; (3) they admit explicit error estimates that only depend on the given characteristic function; (4) multiple values of the desired quantity can be computed simultaneously using the fast Fourier transform. We illustrate the effectiveness of the schemes with financial engineering examples, including the valuation of options with barrier, lookback and early exercise features in Lévy models, as well as Monte Carlo simulation of Lévy processes with analytic characteristic functions.

April 4, 2012

Generalized Inverses Of Matrices: Not Just For Real Or Complex Matrices

Bhaskara Rao Kopparty : 4 p.m. in SEO 636
Abstract Yes. That is true. All of us know that some real and complex matrices have inverses. Not all of them. But we also know that all real and complex matrices have generalized inverses. These are extensively used by statisticians. What about matrices that have only integer entries? Would such a matrix have a generalized inverse whose entries are all integers? What if a matrix has all entries as polynomials? We shall discuss various questions about generalized inverses and see some exciting results. Here is one of the results: An integer matrix of rank r has a generalized inverse if and only if the greatest common divisor of all r x r minors of A is 1.

April 11, 2012

Critical behavior in stochastic models of spatial epidemics

Steven P. Lalley : 4 p.m. in SEO 636
Abstract I will survey of some recent work in scaling limits of stochastic spatial SIR epidemics. In these models, colonies of $N$ individuals are located at lattice points of $\mathbb{Z}^d$. Each individual is, at any time, <i>susceptible</i>, <i>infected</i>, or <i>recovered</i> (and immune to future infection). Infected individuals recover at rate 1, and infect susceptibles in the same or neighboring colonies at rate (say) $\lambda/N$. When $\lambda = 1/(2d + 1)$, where $d$ is the dimension of the lattice, the epidemic is <i>critical</i>: the mean number of new infections produced by a single infected individual when all other individuals in the same or neighboring colonies are susceptible is 1. The questions of natural interest center on the duration and spatial extent of a critical epidemic initiated by a large number $N^\alpha$ (where $0<\alpha <1$) of infected individuals all located at the colony at the origin of the lattice. The main results are large$-N$ <i>scaling laws</i> for the epidemic. The stochastic epidemic models can be reformulated as <i>percolation processes</i> on graphs that are, in a natural sense, hybrids of the standard regular lattices and the complete graph on $N$ vertices. The spread of the epidemic is determined by the geometry of the connected clusters of the associated percolation process. The scaling laws for epidemic processes consequently translate to scaling laws for percolation clusters.

April 18, 2012

Contingent means in multi-life models

Liang Hong : 4 p.m. in SEO 636
Abstract Contingent Means are ubiquitous in modern finance and actuarial science. However, Standard textbooks on actuarial science or statistics do not elaborate on the correct interpretation of contingent means, leaving the actuaries at risk of making a blunder. In this talk, we will give the correct interpretation of contingent means both heuristically and theoretically so that one will be aware of some common misconceptions and avoid pitfalls in their work. We will also discuss the applications of contingent means in insurance and quantitative finance.

April 25, 2012

From a stochastic vortex dynamic model to Onsager-Joyce-Montgomery theory

Jin Feng : 4 p.m. in SEO 636
Abstract The vorticity formulation of 2-D incompressible Navier-Stokes equation can be viewed as mean-field limit of stochastic interacting point vortices. As number of particles goes to infinity and viscosity term goes to zero, we arrive at inviscid limit of 2-D incompressible Euler equation. We study multi-scale large deviation limits of such model, on the torus, as particle number and time go large but viscosity goes small. The result gives a first principle approach to establish the Onsager-Joyce-Montgomery theory as limit theorem derived from microscopically defined non-equilibrium models. The Onsager-Joyce-Montgomery theory concerns large time coherent structures of vortex dynamics associated with 2-D Euler equation. It was previously informally formulated using equilibrium models only. The talk is based on joint works (some of which are ongoing) of the speaker with Fasto Gozzi, Tom Kurtz and Andrzej Swiech, a SQuaRE team funded by American Institute of Mathematics.

June 27, 2012

Outperformance Portfolio Optimization: Hypothesis Testing Approach

Qingshuo Song : 3 p.m. in SEO 636
Abstract We study the portfolio problem of maximizing the out-performance probability over a random benchmark through dynamic trading with a fixed initial capital. Under a general incomplete market framework, this stochastic control problem can be formulated as a composite pure hypothesis testing problem. We analyze the connection between this pure testing problem and its randomized counterpart, and from latter we derive a dual representation for the maximal outperformance probability. Moreover, in a complete market setting, we provide a closed-form solution to the problem of beating a leveraged exchange traded fund. For a general benchmark under an incomplete stochastic factor model, we provide the Hamilton-Jacobi-Bellman PDE characterization for the maximal out-performance probability. It's a joint work with Tim Leung and Jie Yang.

Aug. 29, 2012

Statistics Seminar Organization Meeting

Statistics Group : 4 p.m. in SEO 636

Sept. 5, 2012

Plausibility functions and exact frequentist inference

Prof. Ryan Martin : 4 p.m. in SEO 636
Abstract In the frequentist program, inferential methods with exact control on error rates are a primary focus. Methods based on asymptotic distribution theory may not be suitable in a particular problem, in which case, a numerical method is needed. In this talk I shall present a general, yet simple, Monte Carlo-driven framework for the construction of frequentist procedures based on plausibility functions. It is proved that the suitably defined plausibility function-based tests and confidence regions have desired frequentist properties. Moreover, in an important special case involving likelihood ratios, conditions are given such that the plausibility function behaves asymptotically like a consistent Bayesian posterior distribution. An extension of the proposed method is also given for the case where nuisance parameters are present. I shall give several examples to illustrate the method's flexibility and to demonstrate its performance compared to existing numerical and analytical methods.

Sept. 12, 2012

Patience Pays Off: A Policy Improvement Algorithm for Stochastic Games of Perfect Information with Average Payoffs

Matthew Bourque (PhD Candidate) : 4 p.m. in SEO 636
Abstract Stochastic games model a competitive situation between two players in discrete time steps over an infinite horizon, in which players' payoffs at each stage depend on both players' action choice. They can be seen as generalizations of both repeated games and Markov decision processes (MDPs). Policy improvement algorithms are an important category of fast algorithms for solving MDPs. In this talk, we will give an introduction to stochastic games, in particular zero-sum games of perfect information and with ARAT structure, and discuss a policy improvement algorithm for finding optimal policies for both players for these categories of games when players are evaluating their payoff steams via a limiting average.

Sept. 19, 2012

How to be a Statistical Entrepreneur in a Data Rich Environment

Dr. Stephen G. Eick : 4 p.m. in SEO 636
Abstract The next decade will be an ideal time to become a Statistical Entrepreneur. We are currently at the beginning edge of what will be a massive deployment of sensors, GPS-enabled smartphones, ubiquitous video, and all sorts of smart devices. This trend, often called "The Internet of Things," will involve devices which stream wireless real-time data back into cloud-based servers that will perform analytical calculations. For statisticians, we are entering a data rich environment while there will be tremendous need to create tools to analyze and make sense of this data. After spending nearly 15 years working in academia and for big business, I became a Statistical Entrepreneur. I have been involved with 1⁄2 a dozen emerging growth analytics software companies would like to share some of my experiences. This is by far the best time that I have ever seen for statistical entrepreneurship.

Sept. 26, 2012

Joint high-dimensional Bayesian variable and covariance selection with an application to eQTL analysis

Prof. Anindya Bhadra : 4 p.m. in SEO 636
Abstract We describe a Bayesian technique to (a) perform a sparse joint selection of significant predictor variables and significant inverse covariance matrix elements of the response variables in a high-dimensional linear Gaussian sparse seemingly unrelated regression (SSUR) setting and (b) perform an association analysis between the high-dimensional sets of predictors and responses in such a setting. To search the high-dimensional model space, where both the number of predictors and the number of possibly correlated responses can be larger than the sample size, we demonstrate that a marginalization-based collapsed Gibbs sampler, in combination with spike and slab type of priors, offers a computationally feasible and efficient solution. As an example, we apply our method to an expression quantitative trait loci (eQTL) analysis on publicly available single nucleotide polymorphism (SNP) and gene expression data for humans where the primary interest lies in finding the significant associations between the sets of SNPs and possibly correlated genetic transcripts. Our method also allows for inference on the sparse interaction network of the transcripts (response variables) after accounting for the effect of the SNPs (predictor variables). We exploit properties of Gaussian graphical models to make statements concerning conditional independence of the responses. Our method compares favorably to existing Bayesian approaches developed for this purpose. This is joint work with Bani K. Mallick of Texas A&M University.

Oct. 3, 2012

Covariance and Precision Matrix Estimation for High-Dimensional Time Series

Dr. XiaoHui Chen : 4 p.m. in SEO 636
Abstract Covariance matrix and its inverse (a.k.a. precision matrix) play a central role in a broad range of problems in statistics and machine learning. In the past few years, there has been an explosion of interest in regularized covariance and precision matrix estimation for high-dimensional i.i.d. random vectors with sub-Gaussian tails. In this talk, we shall discuss the estimation of covariance and precision matrices for stationary and locally stationary high-dimensional time series. In the latter case, the covariance matrices evolve smoothly in time and thus form a covariance matrix function. Under the framework of Wu (2005)'s functional dependence measure, we obtain the rate of convergence for the thresholded covariance matrix estimate and illustrate how the dependence affects the rate of convergence. Asymptotic properties are also obtained for the precision matrix estimate based on the graphical Lasso principle. Our theory substantially generalizes earlier ones by allowing dependence, by allowing non-stationarity and by relaxing the associated moment conditions. Our new results have implications on a number of classical problems, including spatial-temporal statistics and graphical models, among many others.

Oct. 10, 2012

BIG Statistics

Prof. Dennis Lin : 4 p.m. in SEO 636
Abstract In the past decades, we have witnessed the revolution of information technology. Its impact to statistical research is enormous. This talk attempts to address some recent developments and potential research issues in Business, Industry and Government (BIG) Statistics, with special focus on computer experiment and information systems. An overall introduction and review will be given, followed by specific research potentials. Some initial results will be presented, and future research problems will be suggested. If time permits, I will also discuss some recent advances in Search Engine and RFID study. Slides of this talk can be downloaded at the website http://www.personal.psu.edu/users/j/x/jxz203/lin/Lin_pub/

Oct. 24, 2012

Asymptotic properties of the estimates of $l_{1}$ penalized regression

Prof. Shuva Gupta : 4 p.m. in SEO 636
Abstract Here we investigate two problems concerning the asymptotic properties of an $l_{1}$ penalized regression. In the first problem we study the asymptotic distribution of the Lasso estimator for regression models with dependent errors. The asymptotic distribution of the Lasso estimator for regression models with independent errors has been investigated by Knight and Fu. Here we extend these results to regression models with a general weak dependence structure. We determine the asymptotic distribution of the Lasso estimator when the number of parameters M is fixed and the number of observations, n, converges to $\infty$. We show that, for an appropriate choice of the tuning parameter of the method, this asymptotic distribution reduces to a multivariate normal distribution. We also provide some illustrative examples. In the second problem, we deal with the asymptotic distribution of residual empirical process of residuals in an adaptive lasso setting. We study the asymptotic properties of the empirical residual process and then use it to investigate the problem of goodness of fit when p<n but increases with n. We explore different applications of residual empirical process eg test of goodness of fit. This work was largely motivated by that of Chen and Lockhart (2001).

Oct. 31, 2012

Generalized Linear and Nonlinear Models for Correlated Data: Overcoming Apparent Limitations in SAS

Dr. Ed Vonesh : 4 p.m. in SEO 636
Abstract Correlated response data, either discrete (nominal, ordinal, counts), continuous or a combination thereof, occur in numerous disciplines and more often than not, require the use of statistical models that are nonlinear in the parameters of interest. Such models include generalized linear and generalized nonlinear models both of which can be further classified according to whether they are marginal or mixed-effects models. In this talk we briefly describe the different types of correlated response data and models encountered in practice. As some of the models can be quite complicated, there will often appear to be certain modeling limitations with available software. For users of SAS, such limitations would appear to include 1) how to conduct likelihood-based inference for nonlinear mixed-effects models with intra-subject correlation; 2) how to fit nonlinear mixed-effects models to data assuming non-Gaussian random effects; and 3) how to fit marginal generalized linear models to correlated response data using second order generalized estimating equations or maximum likelihood estimation. The focus of this talk will be on illustrating how one can fit mixed-effects models in SAS when the random effects are non-Gaussian. Following the work of Nelson et. al. (2006), the approach entails applying probability integral transformations when evaluating an integrated log-likelihood function. This technique is illustrated through an application that requires one to jointly model two dependent variables using a shared non-Gaussian random effect. It requires the user to be familiar with nonlinear mixed-effects models in general and also with how one can fit such models using the SAS procedure NLMIXED.

Nov. 7, 2012

Locally Optimal Designs for Nonlinear Models

Prof. John Stufken : 4 p.m. in SEO 636
Abstract Identifying optimal designs for nonlinear models is a challenging problem for a variety of reasons. While the problem has received considerable attention over the last decades, a new approach has facilitated significant progress during the last three years in the context of local optimality. Although locally optimal designs are not always the preferred choice for applications, they are useful as a benchmark for other designs. In this talk we will cover some of the challenges related to identifying optimal designs for nonlinear models, discuss the new approach for identifying locally optimal designs that is based on finding small complete classes of designs, and illustrate the power of the approach through examples.

Nov. 14, 2012

Bayesian Empirical Likelihood for Quantile Regression

Prof. Xuming He : 4 p.m. in SEO 636
Abstract Quantile regression is semiparametric in the sense that no parametric likelihood is assumed in the model. A working likelihood can be used, but the resulting posterior may not have the validity for statistical inference. In this talk we will introduce Bayesian empirical likelihood for quantile regression, and show that it leads to asymptotically valid posterior inference. In addition, this approach enables us to make use of commonality across quantiles to improve efficiency of quantile estimation in data sparse areas. We will also introduce a notion of shrinking priors, and demonstrate how this new framework can help explain the efficiency gains of the Bayesian empirical likelihood method over the usual quantile estimates. The talk is based on joint work with Yunwen Yang (Drexel University).

Nov. 21, 2012

Fractal Properties of Stochastic Heat Equations

Prof Yimin Xiao : 4 p.m. in SEO 636
Abstract By applying methods for studying Gaussian random fields, we investigate various analytic and fractal properties of the solutions of the stochastic heat equations driven by space-time white noise or fractional colored noise. These include fractal dimensions, exact modulus of continuity, hitting probabilities and existence of intersections. The proofs of these results are based on the properties of strong local nondeterminism. (Based on the on-going joint works with R. Dalang, C. Mueller and C. Tudor.)

Nov. 28, 2012

Asymptotic Bayes optimality under sparsity of multiple testing and model selection rules

Prof. Malgorzata Bogdan : 4 p.m. in SEO 636
Abstract In Bogdan et al (Ann. Statist. 2011) the asymptotic framework for the analysis of the Bayes risk of the multiple testing procedures under sparsity is proposed. Within this framework the rule is called Asymptotically Bayes Optimal under Sparsity (ABOS) if the ratio of its risk and the risk of the Bayes oracle converges to 1 as the number of tests, $m$, diverges to infinity and the proportion of alternatives among all tests, $p$, converges to zero. In Bogdan et al (2011) and Neuvial and Roquain (Ann. Statist., to appear) the conditions under which the popular Benjamini-Hochberg and Bonferroni procedures are ABOS are provided. We will discuss these results and provide an extension to the situation where the sample size $n$ used to calculate each of the test statistics goes to infinity with the number of tests $m$. We show that under mild restrictions on the loss function and the distribution of the magnitude of true signals a nontrivial asymptotic inference is possible only if $n$ increases to infinity at least at the rate of $\log m$. Based on this assumption precise conditions are given under which the Bonferroni correction with nominal Family Wise Error Rate (FWER) level $\alpha$ and the Benjamini-Hochberg procedure (BH) at FDR level $\alpha$ are asymptotically optimal. In the second part of this talk these optimality results are carried over to model selection in the context of multiple regression with orthogonal regressors. Several modifications of Bayesian Information Criterion are considered, controlling either FWER or FDR, and conditions are provided under which these selection criteria are ABOS. Finally the performance of the multiple testing rules and the model selection criteria is examined in a brief simulation study.

Dec. 5, 2012

Penalized Quantile Regression for Ultra-high Dimensional Data

Prof. Runze Li : 4 p.m. in SEO 636
Abstract Ultra-high dimensional data often display heterogeneity due to either heteroscedastic variance or other forms of non-location-scale covariate effects. To accommodate heterogeneity, we advocate a more general interpretation of sparsity which assumes that only a small number of covariates influence the conditional distribution of the response variable given all candidate covariates; however, the sets of relevant covariates may differ when we consider different segments of the conditional distribution. In this talk, I first introduce recent development on the methodology and theory of nonconvex penalized quantile linear regression in ultra-high dimension. I further propose a two-stage feature screening and cleaning procedure to study the estimation of the index parameter in heteroscedastic single-index models with ultrahigh dimensional covariates. Sampling properties of the proposed procedures are studied. Finite sample performance of the proposed procedure is examined by Monte Carlo simulation studies. A real example example is used to illustrate the proposed methodology.

Jan. 23, 2013

On a Scale Invariant Model of Statistical Mechanics, Kinetic Theory of Ideal Gas, and Riemann Hypothesis

Prof. SIAVASH H. SOHRAB : 4:15 p.m. in SEO 636
Abstract A scale invariant model of statistical mechanics is applied to derive invariant forms of conservation equations. A modified form of Cauchy stress tensor for fluid is presented that leads to modified Stokes assumption thus a finite coefficient of bulk viscosity. The phenomenon of Brownian motion is described as the state of equilibrium between suspended particles and molecular clusters that themselves possess Brownian motion. Physical space or Casimir vacuum is identified as a tachyonic fluid that is "stochastic ether" of Dirac or "hidden thermostat" of de Broglie, and is compressible in accordance with Planck's compressible ether. The stochastic definitions of Planck h and Boltzmann k constants are shown to respectively relate to the spatial and the temporal aspects of vacuum fluctuations. Hence, a modified definition of thermodynamic temperature is introduced that leads to predicted velocity of sound in agreement with observations. Also, a modified value of Joule-Mayer mechanical equivalent of heat is identified as the universal gas constant and is called De Pretto number 8338 which occurred in his mass-energy equivalence equation. Applying Boltzmann's combinatoric methods, invariant forms of Boltzmann, Planck, and Maxwell-Boltzmann distribution functions for equilibrium statistical fields including that of isotropic stationary turbulence are derived. The latter is shown to lead to the definitions of (electron, photon, neutrino) as the mostprobable equilibrium sizes of (photon, neutrino, tachyon) clusters, respectively. The physical basis for the coincidence of normalized spacings between zeros of Riemann zeta function and the normalized Maxwell-Boltzmann distribution and its connections to Riemann Hypothesis are examined. The zeros of Riemann zeta function are related to the zeros of particle velocities or "stationary states" through Euler's golden key thus providing a physical explanation for the location of the critical line. It is argued that because the energy spectrum of Casimir vacuum will be governed by Schrödinger equation of quantum mechanics, in view of Heisenberg matrix mechanics physical space should be described by noncommutative spectral geometry of Connes. Invariant forms of transport coefficients suggesting finite values of gravitational viscosity as well as hierarchies of vacua and absolute zero temperatures are described. Some of the implications of the results to the problem of thermodynamic irreversibility and Poincaré recurrence theorem are addressed. Invariant modified form of the first law of thermodynamics is derived and a modified definition of entropy is introduced that closes the gap between radiation and gas theory. Finally, new paradigms for hydrodynamic foundations of both Schrödinger as well as Dirac wave equations and transitions between Bohr stationary states in quantum mechanics are discussed.

Feb. 6, 2013

Bayesian estimation of sparse high-dimensional normal means

Prof. Ryan Martin : 4 p.m. in SEO 636
Abstract In high-dimensional problems, the parameter of interest is called sparse if most of its components are zero. An important example is in high-dimensional regression, were it is believed that only a few of the many predictor variables explain variation in the response. Since we don't know beforehand which coordinates of the parameter vector are zero, the challenge is to simultaneously identify those which are zero and accurately estimate those which are non-zero. From a Bayesian point of view, to accommodate sparsity, an intuitive strategy is to consider a discrete-continuous mixture prior which allows coordinates to be exactly zero with positive probability. The relevant asymptotic theory looks for conditions such that the posterior distribution concentrates around the true signal at the best rate. In the first part of the talk, I will discuss some results along these lines from the very-recent literature. One drawback to the discrete-continuous mixture priors is that computation can be very difficult, so it is natural to ask if similar posterior concentration results can be achieve with computationally simpler non-mixture priors. This question is almost completely open, and the second part of the talk will discuss some aspects of this problem and what I think can be done.

Feb. 13, 2013

Sharp estimates on the Dirichlet heat kernels of subordinate Brownian motions

Prof Song Renming : 4 p.m. in SEO 636
Abstract A subordinate Brownian motion can be obtained by replacing the time parameter of a Brownian motion by an increasing Levy process (i.e., subordinator). Subordinate Brownian motions are very important in various applications. In this talk, I will give a survey of some recent results in the study of subordinate Brownian motions. In particular, I will present results on sharp two-sided estimates of the Dirichlet heat kernel estimates of subordinate Brownian motions.

Feb. 20, 2013

Statistical Applications in the Non-Clinical Field

Gerald (Jerry) Phillips : 4 p.m. in SEO 636
Abstract The non-clinical field in the healthcare industry is an important area that primarily consists of R&D, manufacturing and product field maintenance. Statistical applications in these areas will be demonstrated. In addition, skills required to provide successful statistical consulting will be discussed.

Feb. 27, 2013

Dihedral Fourier Analysis - Data-Analytic Aspects

Prof. Marlos Viana : 4 p.m. in SEO 636
Abstract The seminar will present an overview of Fourier analysis over the dihedral groups as a method of determining (dihedral) orbit invariants as statistical summaries of (dihedral) experiments. While classical commutative harmonic analysis has a long history in science in general and in optics and vision studies in particular, the dihedral groups are certainly among the finite non-commutative groups of broadest application in those areas. The data-analytic aspects to be discussed include the notions of dihedral experiments, labeling arbitrariness of group orbits, determination and interpretation of the Fourier transforms as orbit invariants and statistical summaries. Some of the applications to briefly illustrated include the modeling of corneal curvature and power surfaces, polarimetric-enhanced retinal imaging methods, dihedral polynomial methods for wave-front aberration analysis, visual field decompositions, and visual perception studies. The theory and methods of dihedral analysis are also applicable in the studies of symbolic sequences in structural/functional molecular biology, vibrational spectroscopy, and several other fields, thus leading to a rich interplay of several basic disciplines such as physics, biology, algebra and statistics among others.

March 6, 2013

Large Portfolio Allocation Using High-Frequency Financial Data

Prof. Jian Zou : 4 p.m. in SEO 636
Abstract Portfolio allocation is one of the most fundamental problems in finance. The process of determining the optimal mix of assets to hold in the portfolio is a very important issue in risk management. It involves dividing an investment portfolio among different assets based on the volatilities of the asset returns. In the recent decades, it gains popularity to estimate volatilities of asset returns based on high-frequency data in financial economics. However the most available methods are not directly applicable when the number of assets involved is large, since small component-wise estimation errors could accumulate to large matrix-wise errors. This paper starts with a review on portfolio allocation and high-frequency financial time series. Then we introduce a new methodology to carry out efficient asset allocations using regularization on estimated integrated volatility via intra-day high-frequency data. We illustrate the methodology with the high-frequency price data on stocks traded in New York Stock Exchange over a period of 209 days in 2010. The theory and numerical results show that our approach perform well in portfolio allocation by pooling together the strengths of regularization and estimation from a high-frequency finance perspective.

March 13, 2013

Gaussian bounds and hitting probabilities for differential equations driven by a fractional Brownian motion

Prof. Fabrice Baudoin : 4 p.m. in SEO 636
Abstract This talk investigates several properties related to densities of solutions $(X_t)_{t\in[0,1]}$ to differential equations driven by a fractional Brownian motion with Hurst parameter $H>1/4$. We first determine conditions for strict positivity of the density of $X_t$. Then we obtain some exponential bounds for this density when the diffusion coefficient satisfies an elliptic type condition. Finally, still in the elliptic case, we derive some bounds on the hitting probabilities of sets by fractional differential systems in terms of Newtonian capacities.

March 20, 2013

Introduction of Capital One Statistician Job Family

Dr. Jennifer Hill : 4 p.m. in SEO 636

April 3, 2013

Functional Principal Component Analysis of Spatial-Temporal Point Processes with Applications in Disease Surveillance

Prof. Yehua Li : 4 p.m. in SEO 636
Abstract In disease surveillance applications, the disease events are modeled by spatial-temporal point processes. We propose a new class of semiparametric generalized linear mixed Cox model for such data, where the event rate is related to some known risk factors and some unknown latent random effects. We model the latent spatial-temporal process as spatially correlated functional data, and propose composite likelihood methods based on spline approximation to estimate the mean and covariance of the latent process. By performing functional principal component analysis to the latent process, we gain deeper understanding of the correlation structure in the point process, and we propose an empirical Bayes method to predict the latent spatial random effects, which can help highlighting the high risk spatial regions for the disease. Under an increasing domain and increasing knots asymptotic framework, we provide the asymptotic distribution for the parametric components in the model and the asymptotic convergence rate for the functional principal component estimators. We illustrate the methodology through a simulation study and an application to the Connecticut Tumor Registry data.

April 10, 2013

Demystifying Gaussian Process Models with Massive Data

Prof. Peter Qian : 4 p.m. in SEO 636
Abstract Gaussian process (GP) models are widely used in statistics, optimization, machine learning and other fields. Fitting a GP model with massive data is not only a challenge but also a mystery. On one hand, the nominal accuracy of a GP model is supposed to increase with the number of data points. On the other hand, fitting such a model to a large number of points encounters numerical singularity. To reconcile this contradiction, I will present a method to achieve both numerical stability and theoretical accuracy in fitting a massive GP model. This method obtains nested subsamples of the data, builds submodels for different subsets and then combines these models together to form an accurate prediction model. A decomposition of the overall model error into nominal and numeric portions is introduced to shed light on the theoretical underpinnings of the method. Bounds on the numeric and nominal error are developed to show that substantial gains in overall accuracy can be attained with this sequential method. Efficient algorithms are introduced to generate the required nested subsamples of the developed method.

April 17, 2013

Bayesian problems with normalizing constants

Prof. Stephen Walker : 4 p.m. in SEO 636
Abstract There are desirable models with attractive features which suffer from the problem of intractable normalizing constants. A general strategy for dealing with such constants in a Bayesian setting is presented. A range of illustrations, from the Fisher-Bingham distribution to power likelihood models, regression and time series models will be provided.

April 24, 2013

Brownian Motion on Spaces with Varying Dimension

Prof. Zhen-Qing Chen : 4 p.m. in SEO 636
Abstract Image an insect moves randomly in a plane with an infinite pole installed on it. In this talk, we will introduce and discuss Brownian motion on a state space with varying dimension. We will derive sharp two-sided estimates on its transition density function (also called heat kernel). The two-sided estimates is of Guassian type but the parabolic Harnack inequality fails for such process and the measure on the underlying state space does not satisfy volume doubling property.

May 1, 2013

Inference Function for Mixed Effects Models and its applications

Prof. Peng Wang : 4 p.m. in SEO 636
Abstract In longitudinal studies, mixed-effects models are important for addressing subject-specific effects. However, most existing approaches assume a normal distribution for the random effects, and this could affect the bias and efficiency of the fixed-effects estimator. Even in cases where the estimation of the fixed effects is robust with a misspecified distribution of the random effects, the estimation of the random effects could be invalid. We propose a new approach to estimate fixed and random effects using conditional quadratic inference functions. The new approach does not require the specification of likelihood functions or a normality assumption for random effects. It can also accommodate serial correlation between observations within the same cluster, in addition to mixed-effects modeling. Other advantages include not requiring the estimation of the unknown variance components associated with the random effects, or the nuisance parameters associated with the working correlations. We establish asymptotic results for the fixed-effect parameter estimators which do not rely on the consistency of the random-effect estimators. Some applications of the proposed approach will also be presented.

Aug. 28, 2013

Statistics group gathering

Statistics group : 4 p.m. in SEO 636
Abstract Welcome new students and discuss and expose students to various statistical associations. This gathering is partially sponsored by the American Statistical Association.

Sept. 4, 2013

Overview of Agreement Statistics for Continuous, Binary, and Ordinal Data

Dr. Lawrence Lin : 4 p.m. in SEO 636
Abstract This will be a general overview presentation with practical examples and without much statistical formulas. We will introduce the concepts of un-scaled and scaled agreement statistics based on the basic case between two raters with paired samples for continuous, binary, and ordinal data. We will then progress into more complex cases when we have multiple raters and each rater has multiple readings per sample. Here, we can assess intra-rater and inter-rater agreement, compare inter-rater deviation to intra-rater deviation, and compare precision of a rater against another. We will explore the meaning of the two-stage criteria presented in the FDA guidance UCM070244: Statistical Approaches to Establishing Bioequivalence. The content is largely based on the materials presented in the newly published book by Springer, entitled "Statistical Tools for Assessing Agreement".

Sept. 11, 2013

Statistical Methods for Analysis of Graph constrained Genomic data

Dr. Caiyan Li : 4 p.m. in SEO 636
Abstract Graphs and networks are common ways of depicting biological information. In biology, many different biological processes are represented by graphs, such as regulatory networks, metabolic pathways and protein protein interaction networks. This kind of a priori use of graphs is a useful supplement to the standard numerical data such as microarray gene expression data. In this presentation, we consider the problem of regression analysis and variable selection when the covariates are linked on a graph. We study a graph constrained regularization procedure and its theoretical properties for regression analysis to take into account the neighborhood information of the variables measured on a graph. This procedure involves a smoothness penalty on the coefficients that is defined as a quadratic form of the Laplacian matrix associated with the graph. We establish estimation and model selection consistency results and provide estimation bounds for both fixed and diverging numbers of parameters in regression models. We also developed a second method using Markov Random Field to incorporate the graph information into analysis of high dimensional data. Finally, we demonstrate by simulations and a real dataset that the proposed procedure can lead to better variable selection and prediction than existing methods that ignore the graph information associated with the covariates.

Sept. 18, 2013

The optimal weights exchange algorithm – A fast and flexible method to derive optimal designs

Dr. Stefanie Biedermann : 4 p.m. in SEO 636
Abstract Finding optimal designs for nonlinear models is challenging in general. Although some recent results allow us to focus on a simple subclass of designs for most problems, deriving a specific optimal design still mainly depends on numerical approaches. There is need for a general and efficient algorithm which is more broadly applicable than the current state of the art methods. We present a new algorithm which can be used to find optimal designs with respect to a broad class of optimality criteria, when the model parameters or functions thereof are of interest, and for both locally optimal and multi-stage design strategies. We prove convergence to the optimal design, and show in various examples that the new algorithm outperforms the current state of the art algorithms.

Sept. 25, 2013

Detection of Disease Outbreaks with Complex Spatio-Temporal Structures Using Real Surveillance Data

Prof. Jian (Frank) Zou : 4 p.m. in SEO 636
Abstract The complexity of spatio-temporal data in epidemiology and surveillance presents challenges such as low signal-to-noise ratio and generating high false positive rate for researchers and public health agencies. Central to the problem in the context of disease outbreaks is a decision structure that requires trading off false positives for delayed detections. We describe a novel Bayesian hierarchical model capturing the spatio-temporal dynamics in public health surveillance data sets. We further quantify the performance of the method to detect outbreaks by incorporating different criteria, including false alarm rate, timeliness and cost functions. Our data set is derived from emergency department (ED) visits for Influenza-like illness and respiratory illness in the Indiana Public Health Emergency Surveillance System (PHESS). The methodology incorporates Gaussian Markov random field (GMRF) and spatio-temporal conditional autoregressive (CAR) modeling. Features of this model include timely detection of outbreaks, robust inference to model misspecification, reasonable prediction performance, as well as attractive analytical and visualization tool to assist public health authorities in risk assessment. Our numerical results show that the model captures salient spatio-temporal dynamics that are present in public health surveillance data sets, and that it appears to detect both "annual" and "atypical" outbreaks in a timely, accurate manner. We present maps that help make model output accessible and comprehensible to public health authorities. We use an illustrative family of decision rules to show how output from the model can be used to inform false positive--delayed detection tradeoffs.

Oct. 2, 2013

Generalized Fiducial Inference and Confidence Distributions

Jan Hannig : 4 p.m. in SEO 636
Abstract R. A. Fisher's fiducial inference has been the subject of many discussions and controversies ever since he introduced the idea during the 1930's. The idea experienced a bumpy ride, to say the least, during its early years and one can safely say that it eventually fell into disfavor among mainstream statisticians. However, it appears to have made a resurgence recently under various names and modifications. For example under the new name generalized inference fiducial inference has proved to be a useful tool for deriving statistical procedures for problems where frequentist methods with good properties were previously unavailable. Therefore we believe that the fiducial argument of R.A. Fisher deserves a fresh look from a new angle. In this talk we first generalize Fisher's fiducial argument and obtain a fiducial recipe applicable in virtually any situation. We demonstrate this fiducial recipe on examples of varying complexity. We also investigate, by simulation and by theoretical considerations, some properties of the statistical procedures derived by the fiducial recipe showing they often posses good repeated sampling, frequentist properties. Finally, we show how a generalized fiducial inference paradigm can be used to combine confidence distributions. Portions of this talk are based on a joined work with Hari Iyer, Thomas C.M. Lee, Randy Lai and Min-ge Xie.

Oct. 9, 2013

Asymptotically minimax empirical Bayes estimation of a sparse normal mean vector

Ryan Martin : 4 p.m. in SEO 636
Abstract Estimating a sparse high-dimensional normal mean vector is an important classical problem. In this talk, I will introduce a new empirical Bayes model based on a unique data-dependent prior. I will show that, under some conditions, our empirical Bayes posterior distribution concentrates on balls, centered at the true mean vector, with squared radius proportional to the frequentist minimax rate for the given sparsity class. This result provides some new insight concerning the fully Bayes approach to this same problem. Asymptotic minimaxity of the corresponding empirical Bayes posterior mean is shown, and a simple Gibbs sampling algorithm for computation will be discussed. Finally, two simulation studies will be presented, demonstrating the strong finite-sample performance of the proposed estimator against a variety of popular alternatives. (This is joint work with Stephen Walker at the University of Texas at Austin.)

Oct. 16, 2013

Templates for Design Key Construction

Prof. Ching-Shui Cheng : 4 p.m. in SEO 636
Abstract Design key is a useful method of constructing factorial designs with what J. A. Nelder called simple block structures. One familiar method of constructing such designs is to use independent treatment factorial effects to divide the treatment combinations into blocks, rows, and columns, etc. Certain conditions need to be verified to ensure that the desired block structure is achieved. The method of design key, proposed by H. D. Patterson, on the other hand, guarantees the desired block structure by choosing some appropriate contrasts of the experimental units to be aliases of the main-effect contrasts of treatment factors. We provide some useful templates for implementing design key construction. Some constraints imposed by the block structure are built into the template. This eliminates the need to check some conditions for design eligibility.

Oct. 23, 2013

Malliavin calculus and convergence in density of some nonlinear Gaussian functionals

Prof. Yaozhong Hu : 4 p.m. in SEO 636
Abstract The classical central limit theorem is one of the most important theorem in probability theory. The theorem states that if $X_1$, $\cdots$, $X_n$ are independent identically distributed random variables and if $F_n$ is the difference between the sample mean and the mean of the random variables properly normalized, then $F_n$ converges to a normal distribution in distribution. Recent results extend this results to other random variables for example given by Wiener chaos (multiple It\^o-Wiener integrals). In this talk, we shall obtain some conditions on $F_n$ such that the distributions of the random variables $F_n$ have densities $f_n(x)$ with respect to Lebesgue measure and $f_n(x)$ converges to the normal density $\phi(x)=\frac{1}{\sqrt{2\pi}}e^{-|x|^2/2}$. The tool that we use is the Malliavin calculus and a brief introduction will also be given. This is an ongoing joint work with Fei Lu and David Nualart

Oct. 30, 2013

Working as a Statistician in Pharmaceutical Industry

Ms. Ping Jiang : 4 p.m. in SEO 636
Abstract Working as a statistician in a pharmaceutical company is exciting and rewarding. This presentation will show the vital roles played by statisticians in various stages of drug research and development. The information is intended for the students who are interested in pursuing a career in pharmaceutical industry.

Nov. 6, 2013

Nonparametric Regression With Errors in Partial Variables

Prof. Weixing Song : 4 p.m. in SEO 636
Abstract An estimation procedure is proposed for a nonparametric regression in which some covariates are measured with errors and some are not. The procedure combines the ordinary and deconvolution kernel estimation techniques. It is shown that the optimal local and global convergence rates, as well as the uniform convergence rate over a class of joint distributions of the response and the covariates, depend on the tail behavior of the characteristic functions of the measurement error distributions. Examples are given to show the general applicability of the proposed methodology, and finite sample performance is evaluated by some numerical simulation studies.

Nov. 13, 2013

Jump detection in time series nonparametric regression models: a polynomial spline approach

Prof. Qiongxia Song : 4 p.m. in SEO 636
Abstract For time series nonparametric regression models with discontinuities, we propose to use polynomial splines to estimate locations and sizes of jumps in the mean function. Under reasonable conditions, test statistics for the existence of jumps are given and their limiting distributions are derived under the null hypothesis that the mean function is smooth. Simulations are provided to check the powers of the tests. A climate data application and an application to the U.S. unemployment rates of men and women are used to illustrate the performance of the proposed method in practice.

Nov. 20, 2013

Exchangeable Markov process models for time-varying networks

Prof. Harry Crane : 4 p.m. in SEO 636
Abstract In fields as diverse as physics, biology, sociology and national security, complex networks are used to model structural relationships among individuals and variables. In many applications, the networks vary over time and so are appropriately modeled by a stochastic process on the space of graphs. Motivated by these applications, we consider Markov processes that evolve on the space of infinite graphs. Natural statistical models for such processes are both exchangeable with respect to relabeling vertices and have the property that all restrictions to finite induced subgraphs are finite state space Markov chains. Our main theorem provides a Levy-Ito-type characterization for all processes in this class. Our approach also gives a straightforward recipe for simulating general processes of this type, which may be useful in a range of applications.

Jan. 22, 2014

Brownian Motion with Varying Dimension

Shuwen Lou : 4 p.m. in SEO 636
Abstract Think of installing an infinite pole on top of the ground. We want to model the random movement of an ant on this space. However, as we know, the standard 2-dimensional Brownian motion does not hit a single, which means that once the ant is on the ground, it will never have the chance to climb up the pole. We fix this problem by defining Brownian motion with varying dimension on this state space as a darning process whose rigorous definition will be introduced in the talk. The main results are about global two-sided heat kernel estimates for such processes. We will see from the heat kernel estimates that these processes embody both 1-dimensional property and 2-dimensional property, which depends not only on the regions of the points but also on time.

Jan. 29, 2014

A Probabilistic Graphical Model for Brand Reputation Assessment in Social Networks

Prof. Kunpeng Zhang : 4 p.m. in SEO 636
Abstract Social media has become a popular platform that connects people who share information, in particular personal opinions. Through such a fast information exchange mechanism, reputation of individuals, consumer products, or business companies can be quickly built up within a social network. Recently, applications mining social network data start emerging to find the communities sharing the same interests for marketing purposes. Knowing the reputation of social network entities, such as celebrities or business companies, can help develop better strategies for election campaigns or new product advertisements. In this work, we propose a probabilistic graphical model to collectively measure reputations of entities in social networks. By collecting and analyzing large amount of user activities on Facebook, our model can effectively and efficiently rank entities, such as presidential candidates, professional sport teams, musician bands, and companies, based on their social reputation. The proposed model produces results largely consistent with the two publicly available systems - movie ranking in Internet Movie Database and business school ranking by the US news & World Report - with the correlation coefficients of 0.75 and −0.71, respectively. In addition, I will briefly talk about other projects I am working on: (1) sentiment identification of social media data, and (2) finding target users for online brand advertising based on a large amount of user historical activities.

Feb. 12, 2014

Quantum Computation and Statistics

Prof. Yazhen Wang : 4 p.m. in SEO 636
Abstract Quantum computation and quantum information are of great current interest in computer science, mathematics, physical sciences and engineering. They will likely lead to a new wave of technological innovations in communication, computation and cryptography. As the theory of quantum physics is fundamentally stochastic, randomness and uncertainty are deeply rooted in quantum computation, quantum simulation and quantum information. Consequently quantum algorithms are random in nature, and quantum simulation utilizes Monte Carlo techniques extensively. Thus statistics can play an important role in quantum computation and quantum simulation, which in turn offer great potential to revolutionize computational statistics. This talk will give a brief review on quantum computation, quantum simulation and quantum information. I will first introduce the basic concepts of quantum computation and quantum simulation and then present my recent work on statistical analysis of quantum systems with applications to quantum computation and quantum simulation.

March 12, 2014

Personalized treatment for longitudinal data

Prof. Annie Qu : 4 p.m. in SEO 636
Abstract We develop new modeling and estimation for personalized treatment for individuals with high heterogeneity. Incorporating subject-specific information into treatment subgroup is critical since individuals could react to the same treatment quite differently. We propose to identify subgroups with longitudinal observations through random-effects estimation where the random effects are not necessarily normal distributed. The advantage of this approach is that we can quantify intrinsic associations between unobserved subject-specific effects and observed treatment outcomes, and therefore provide optimal treatment assignments for different individuals. In contrast, traditional mixed-effects models assuming normal distribution cannot effectively distinguish different patterns of treatment effects. We develop asymptotic consistency theory for individual treatment effect estimation, and show that the new estimator is more efficient than the random effect estimator which ignores correlation information from longitudinal data. Simulation studies and a data example from an AIDS clinical trial group confirm that the proposed method is quite efficient in identifying an effective treatment strategy for subgroups in finite samples. This is joint work with Hyunkeun Cho and Peng Wang.

March 19, 2014

Predicting Right Score Distributions from Data Collected under Formula Score Instructions

Dr. Hongwen Guo : 4 p.m. in SEO 636
Abstract Under formula score instruction (FSI), test takers may omit items instead of guessing. If the students are required to answer every items (under the rights only scoring instruction, ROI), the score distribution will be different. In this study, we tried to model the data using a simple statistical model and provide a formula to predict the score distribution under ROI based on the score distribution and the omit rate observed under FSI. A preliminary investigation of the guessing parameter is presented in the paper. Based on the data used in the study, the guessing parameter may be close or slightly lower than the chance score.

April 2, 2014

Estimation of the Number of Non-identifiable Species

Prof. Bikas Sinha : 4 p.m. in SEO 636
Abstract Noted environmental statistician Professor Anil Gore, University of Pune, dealt with the problem of estimation of non-identifiable bird species in BNHS Survey. This study was conducted around mid 1970's. In recent time, Dr. Tommy Wright of US Census Bureau posed a similar problem. The speaker had the opportunity to be involved in both the studies in a limited way. He will discuss salient features of the problem and the proposed solutions.

April 9, 2014

Statistical issues with drawing associations between rare genetic variants and common human diseases.

Paul Livermore Auer : 4 p.m. in SEO 636
Abstract To date, much of the work in understanding the genetic basis of common diseases has focused on genetic variants that exist with high frequency in human populations. Recently, the focus has shifted to investigating the role that rare genetic variants may play. Genetic studies of rare variants bring with them a host of statistical issues including decreased statistical power and missing data. We examine these issues and demonstrate the performance of solutions with both real and simulated data.

April 16, 2014

Recent developments in optimal experimental designs for functional MRI

Prof. Ming-Hung (Jason) Kao : 4 p.m. in SEO 636
Abstract Functional magnetic resonance imaging (fMRI) is one of the leading brain mapping technologies for studying brain activity in response to mental stimuli. For neuroimaging studies utilizing this pioneering technology, there is a great demand of high-quality experimental designs that help to collect informative data to make precise and valid inference about brain functions. In this talk, I briefly introduce some recently developed analytical and computational results on fMRI experimental designs. The performance of some commonly considered designs such as m-sequences is discussed. In addition, a new type of fMRI designs that are constructed using a certain type of Hadamard matrices is also discussed. Under certain assumptions, these designs can be shown to be optimal in some statistically meaningful sense. Some possible future research directions are also presented.

April 30, 2014

Econometric analysis of multivariate realised QML: estimation of the covariation of equity prices under asynchronous trading

Prof. Dacheng Xiu : 4 p.m. in SEO 636
Abstract Estimating the covariance between assets using high frequency data is challenging due to market microstructure effects and asynchronous trading. In this paper we develop a multivariate realised quasi-likelihood (QML) approach, carrying out inference as if the observations arise from an asynchronously observed vector scaled Brownian model observed with error. Under stochastic volatility the resulting realised QML estimator is positive semi-definite, uses all available data, is consistent and asymptotically mixed normal. The quasi-likelihood is computed using a Kalman filter and optimised using a relatively simple EM algorithm which scales well with the number of assets. We derive the theoretical properties of the estimator and prove that it achieves the efficient rate of convergence. The estimator is also analysed using Monte Carlo methods and applied to equity data with varying levels of liquidity.

Aug. 13, 2014

Small area estimation with uncertain random effects

Abhyuday Mandal : 11 a.m. in SEO 612
Abstract Random effects models play an important role in model-based small area estimation. Random effects account for any lack of fit of a regression model for the population means of small areas on a set of explanatory variables. In a recent paper, Datta, Hall and Mandal (2011, J. Amer. Statist. Assoc.) showed that if the random effects to account for a lack of fit of a regression model can be dispensed with through a statistical test, then the model parameters and the small area means can be estimated with substantially higher accuracy. The work of Datta et al. (2011) is most useful when the number of small areas, m, is moderately large. For large m, the null hypothesis of no random effects will likely be rejected. Rejection of null hypothesis is usually caused by a few large residuals signifying a departure of the direct estimator (Yi) from the synthetic regression estimator. As a flexible alternative to the Fay-Herriot random effects model and the approach in Datta et al. (2011), in this paper we consider a mixture model for random effects. It is reasonably expected that small areas with population means explained adequately by covariates have little model error, and the other areas with means not adequately explained by covariates will require a random component added to the regression model. This model is a flexible alternative to the usual random effects model and the data determine the extent of lack of fit of the regression model for a particular small area, and include a random effect if needed. Unlike the Datta et al. (2011) approach which recommends excluding random effects from all small areas if a test of null hypothesis of no random effects is not rejected, the present model is less restrictive. We used this mixture model to estimate poverty ratios for 5- to 17-year old related children for the 50 U.S. states and Washington, DC. This application is motivated by the SAIPE project of the US Census Bureau. We empirically evaluated the accuracy of the direct estimates and the estimates obtained from our mixture model and the Fay-Herriot random effects model. These empirical evaluations and a simulation study, in conjunction with a measure of uncertainty of the new estimates show that they are more accurate than the frequentist and the Bayes estimates resulting from the standard Fay-Herriot model.

Aug. 27, 2014

Organizational Meeting

Statistics Faculty : 4 p.m. in SEO 636

Sept. 3, 2014

Stat Wars Episode VI: Return of the Fiducialist

Keli Liu : 4 p.m. in SEO 636
Abstract Priors are the path to the dark side. Fisher developed the Fiducial argument to obtain prior free "posterior" inferences but at the seeming cost of violating basic probability laws. Was Fisher crazy or did madness mask innovation? Fiducial calculations can be easily understood through the missing-data perspective which illuminates for us that the Fiducial "posterior" is in fact a prior updated not with the full data likelihood, but a <i>partial</i> likelihood in the spirit of Cox regression. Just as Cox regression arose from a need to render inferences robust to an unknown hazard function, so Fiducial inferences are insensitive to the prior. While Statistics has fixated two extremes---fully conditional (but fragile) Bayesian inferences or unconditional (but robust) Frequentist inferences---a compromise via partial conditioning has gone ignored. Surely, the middle ground is more fiducial than either extreme.

Sept. 17, 2014

Estimating the size of a population comprising of individuals with indirect accessibility

Bikas Sinha : 4 p.m. in SEO 636
Abstract Considered is the problem of estimation of the size of a finite labeled population. However, the units of the population are not directly accessible. This is a well defined finite population of Reference Units (RUs) and these RUs may be accessed directly. As against this, the former population is referred to as that of Ultimate Units (UUs). Further to this, there is a well defined network connecting the population of RUs to that of the UUs. A sample of RUs will naturally create a partial view of the population network. Based on this sample network, it is required to unbiasedly estimate the size of the population of UUs. We review the literature (which is scanty anyway) and provide some suggestions towards satisfactory acceptable solutions to this fascinating problem.

Sept. 24, 2014

Brownian motion on manifolds

Jennifer Pajda-Delao : 4 p.m. in SEO 636
Abstract I will give some interesting estimates of Brownian motion on manifolds, including exit time estimates from a geodesic ball.

Oct. 1, 2014

Small-time asymptotics and expansions of option prices under Lévy-based models

Ruoting Gong : 4 p.m. in SEO 636
Abstract The small-time asymptotic behavior of option prices and implied volatilities for jump-diffusion models has received much attention in recent years. In this presentation, we study the time-to-maturity asymptotics of call option prices under a variety of models with Lévy jumps. In the out-of-the-money (OTM) and in-the-money (ITM) case, we consider a general stochastic volatility model with independent Lévy jumps for the log-return process of the underlying stock price. In this setting, small-time expansions, of arbitrary polynomial order, in time-t, are obtained for both OTM and ITM call option prices. In the at-the-money (ATM) case, a novel second-order approximation of the call option price is obtained for a large class of exponential "tempered-stable-like" Lévy models with or without Brownian component. As a consequence, small-time expansions of the corresponding Black-Scholes implied volatilities are also addressed in both cases. This is the joint work with J. E. Figueroa-López and C. Houdré.

Oct. 8, 2014

Mock Oral Exam

Stat Grad Students : 4 p.m. in SEO 636

Oct. 15, 2014

L2 asymptotics for high-dimensional data

Mengyu Xu : 4 p.m. in SEO 636
Abstract We develop an asymptotic theory for $L^2$ norms of sample mean vectors of high-dimensional data. An invariance principle for the $L^2$ norms is derived under conditions that involve a delicate interplay between the dimension $p$, the sample size $n$, and the moment condition. Under proper normalization, central and non-central limit theorems are obtained. To facilitate the related statistical inference, we propose a resampling calibration method to approximate the distributions of the $L^2$ norms. Our results are applied to multiple tests and inference of covariance matrix structures.

Oct. 22, 2014

Universally optimal designs for two interference models

Wei Zheng : 4 p.m. in SEO 636
Abstract A systematic study is carried out regarding universally optimal designs under the interference model, previously investigated by Kunert and Martin (2000) and Kunert and Mersmann (2011). Parallel results are also provided for the undirectional interference model, where the left and right neighbor effects are equal. It is further shown that the efficiency of any design under the latter model is at least its efficiency under the former model. Designs universally optimal for both models are also identified. Most importantly, this paper provides Kushner's type linear equations system as a necessary and sufficient condition for a design to be universally optimal. This result is novel for models with at least two sets of treatment-related nuisance parameters, which are left and right neighbor effects here. It sheds light on other models in deriving asymmetric optimal or efficient designs.

Oct. 29, 2014

Some important statistical considerations in biomarker discovery from high-dimensional data

V. Devanarayan : 4 p.m. in SEO 636
Abstract Biomarkers such as those based on genomic, proteomic and imaging modalities play a vital role in biopharmaceutical R&D. Examples include the discovery of novel genes/targets related to various diseases based on which a suitable therapeutic can be developed, diagnostics for different disease subtypes, identification of patients that are more likely to progress in disease or benefit from a particular therapeutic, etc. The discovery of such biomarkers are typically based on the evaluation of high-dimensional datasets that require a strong combination of bioinformatic and statistical considerations. This seminar will provide a practical overview and intuitive explanation of some important concepts and considerations around the analyses of such high-dimensional data.

Nov. 5, 2014

Random graphs and networks: estimation and modeling challenges

Sonja Petrovic : 4 p.m. in SEO 636
Abstract The ubiquity of network data in the world around us does not imply that the statistical modeling and fitting techniques have been able to catch up with the demand. This talk will discuss some of the basic modeling questions that every statistician knows are fundamental, some of the recent advances toward answering them, and the challenges that remain. The specific focus of the talk will be on goodness of fit testing for random graph models. Recent joint work with Despina Stasi and Elizabeth Gross developed a new testing framework for graphs that is based on combinatorics of hypergraphs and model geometry. I will summarize our work by showing simulation results for the popular $p_1$ model for directed random graphs.

Nov. 12, 2014

Optimal Plate Designs in High Throughput Screening Experiments

Xianggui Qu : 4 p.m. in SEO 636
Abstract High-throughput screening (HTS) is a large-scale process that screens hundreds of thousands to millions of compounds in order to identify potentially leading candidates rapidly and accurately. There are many statistically challenging issues in HTS. In this talk, I will focus the spatial effect in primary HTS. I will discuss the consequences of spatial effects in selecting leading compounds and why the current experimental design fails to eliminate these spatial effects. A new class of designs will be proposed for elimination of spatial effects. The new designs have the advantages such as all compounds are comparable within each microplate in spite of the existence of spatial effects; the maximum number of compounds in each microplate is attained, etc. Optimal designs are recommended for HTS experiments with multiple controls.

Nov. 19, 2014

Double empirical Bayes for high-dimensional inference / On Bayesian inference without a model

Raymond Mess / Nick Syring : 4 p.m. in SEO 636
Abstract This is a special graduate student-organized seminar in which two PhD students (Raymond Mess and Nick Syring) will give 20+ minute talks about their ongoing research. The respective abstracts are below. (Mess) In this talk, I will introduce the new double empirical Bayes framework, which is based on the use of data to both center and regularize the prior. An application of this framework to the problem of inference in the sparse (p >> n) linear model will also be presented. (Syring) I will introduce a method to obtain Bayesian-like posterior inference for an unknown parameter without the need for a likelihood. Such a method makes producing interval estimates straightforward while avoiding problems that may arise from model misspecification. Finally, I will discuss an application of this approach to an important problem in medical statistics.

Dec. 3, 2014

Computer Experiment with Qualitative and Quantitative Factors

Dr. Devon Lin : 4 p.m. in SEO 636
Abstract Computer experiments with qualitative and quantitative factors occur frequently in various applications in science and engineering. Design and analysis of such experiments is not yet completely resolved. To address this issue, we propose a new class of designs and a flexible modeling approach. The proposed designs allow us to accommodate a large number of qualitative factors with economic run sizes. Properties of such designs will be discussed. Several construction methods will be given. The new modeling approach employs a flexible function to capture the correlation among qualitative and quantitative factors. Several examples are provided to demonstrate significant improvement in prediction.

Jan. 14, 2015

Stochastic De Giogi iteration and regularity of SPDE

Zhenan Wang : 4 p.m. in SEO 636
Abstract We will start on the classical De Giorgi iteration for parabolic PDEs. We will explain how a stochastic version of De Giorgi iteration can be developed and applied to prove H\"older continuity for solution of stochastic partial differential equations with measurable coefficient. We will also introduce fine properties for the solutions obtained by applying the stochastic De Giorgi iterations.

Jan. 21, 2015

Recursive Bayes prediction with copulas

Ryan Martin : 4 p.m. in SEO 636
Abstract The Bayesian framework provides a nice recipe for constructing the predictive distribution of a future observation given the available data. Except for simple problems, computation of the Bayes predictive requires Monte Carlo which cannot be done recursively. However, when data is received sequentially, e.g., in finance applications, a recursive update to the predictive distribution is desired. In this talk, I will explain how the Bayes predictive step can be rewritten using a copula, which makes recursive updates of the predictive distribution possible. This new representation motivates a version of Newton's predictive recursion algorithm for the predictive density, which can be used for fast and universal recursive predictive density estimation. Illustrations and convergence theory for the new algorithm is provided.

Jan. 28, 2015

Estimates for the Dirichlet heat kernel on inner uniform domains

Janna Lierl : 4 p.m. in SEO 636
Abstract I will present sharp two-sided bounds for the Dirichlet heat kernel on bounded domains. The domain is assumed to satisfy an inner uniformity condition. For example, the interior of the Koch snowflake is an inner uniform domain, as is any convex domain, or the complement of any convex domain in Euclidean space. More generally, we have considered the Dirichlet heat kernel on domains in a metric measure Dirichlet space, assuming the space satisfies a Poincar\'e inquality and has the volume doubling property. We have also considered non-symmetric Dirichlet spaces. In particular, we can estimate the Dirichlet heat kernel if the kernel is associated with a differential operator in divergence form with bounded measurable coefficients and symmetric uniformly elliptic second order part. This talk is based on a joint paper with Laurent Saloff-Coste.

Feb. 11, 2015

A general and efficient algorithm for multiple objective optimal design / Estimation efficiency in continual reassessment method

Qianshun Cheng / Tian Tian : 4 p.m. in SEO 636
Abstract (Cheng) An experiment often has several competing objectives cannot be characterized by only one of the standard optimality criteria. Multiple objective optimal design aims to optimize the target objective while guarantee that efficiency of the other objectives interested are above acceptable levels. Such optimality problem is in general challenging and typically be solved through algorithm approach. The existing approaches either have high computation cost or have low accuracy. In this talk, I will present a new algorithm which can be used for general multiple objective optimal design problems regardless of model settings. Compared with the existing approach, the new algorithm enjoys low computation cost and high accuracy. (Tian) A widely used approach of designing the phase I clinical trial is continual reassessment method (CRM), which has been shown through many simulations to be more effective than other traditional approaches. In this talk, I will show that the CRM algorithm is indeed efficient from the perspective of optimal design theory. Specifically, simple power model and logistic model -- two popular models, are considered. For simple power model, I'll show the efficiency of CRM depends on the target toxicity rate and CRM is highly efficient in practice. A remarkable fact is that the optimal design selects the dose level such that the corresponding toxicity rate is around 0.2, which is exactly the commonly used target toxicity rate in clinical trials. Moreover, by incorporating the idea of optimal design into the study, the percentage of toxicity occurrence in the trial will drop by a great amount. As for logistic model, I'll show that the CRM approach is indeed optimal, which will justify the efficiency of the algorithm in theory.

Feb. 18, 2015

Leveraging Algorithms for Logistic Regression with Massive Data

Haiying Wang : 4 p.m. in SEO 636
Abstract For massive data with super-large sample size n, it is computationally infeasible to obtain maximum likelihood estimates for unknown parameters, especially when the estimator does not have a close-form solution. This paper proposes fast leveraging algorithms to efficiently approximate the maximum likelihood estimates of unknown parameters in logistic regression models with binary responses, one of the most commonly used models in practice for classification. We theoretically prove the consistency of the leveraging algorithms, develop nearly optimal two-step leveraging strategies, and evaluate the performance of the proposed methods using synthetic and real data sets.

Feb. 25, 2015

On the uniqueness and properties of the Parisi measure

Wei-kuo Chen : 4 p.m. in SEO 636
Abstract Spin glasses are disordered spin systems originated from the desire of understanding the strange magnetic behaviors of certain alloys in physics. As mathematical objects, they are often cited as examples of complex systems and have provided several fascinating structures and conjectures. This talk will be focused on one of the most famous mean-field spin glasses, the Sherrington-Kirkpatrick model. We will present results on the conjectured properties of the Parisi measure including its uniqueness and quantitative behaviors. This is based on joint works with A. Auffinger.

March 4, 2015

Model-free variable selection via learning gradients

Lei Yang : 4 p.m. in SEO 636
Abstract Variable selection is popular in high-dimensional data analysis to identify the truly informative variables. Many variable selection methods have been developed under various model assumptions, such as linear model and additive model. However, their success largely rely on validity of the assumed models. In this talk, I will introduce a model-free variable selection method based on gradient learning. The key idea is that if a variable is informative is equivalent to if its corresponding gradient function is substantially non-zero. The proposed method is formulated in a framework of learning gradients equipped with a flexible reproducing kernel Hilbert space. Computationally, a blockwise majorization decent (BMD) algorithm is introduced for efficient computation. Theoretically, without assuming explicit models, the estimation and variable selection consistencies are established. A variety of simulated examples and real-life examples are provided to evaluate the performance.

March 18, 2015

A new nonparametric stationarity test of time series in time domain

Professoe Suojin Wang : 4 p.m. in SEO 636
Abstract In this talk, we present a new double order selection test for checking second-order stationarity of a time series. To develop the test, a sequence of systematic samples are defined via the Walsh functions. Then the deviations of the autocovariances based on these systematic samples from the corresponding autocovariances of the whole time series are calculated and the uniform asymptotic joint normality of these deviations over different systematic samples is obtained. With a double order selection scheme, our test statistic is constructed by combining the deviations at different lags in the systematic samples. The null asymptotic distribution of the proposed statistic is derived and the consistency of the test is shown under fixed and local alternatives. Simulation studies demonstrate well-behaved finite sample properties of the proposed method. Comparisons with some existing tests in terms of power are given both analytically and empirically. In addition, the proposed method is applied to check the stationarity assumption of a chemical process viscosity readings data.

April 8, 2015

Statistical properties of eigenvalues of Laplace-Beltrami operators

Tiefeng Jiang : 4 p.m. in SEO 636
Abstract We study the eigenvalues of a Laplace-Beltrami operator defined on the set of the symmetric polynomials, where the eigenvalues are expressed in terms of partitions of integers. By assigning partitions with the uniform measure, the restricted uniform measure, the Plancherel measure, or the restricted Jack measure, we prove that the global distributions of the eigenvalues are asymptotically the Gumbel distribution, a new distribution F, the Tracy-Widom distribution and the Gamma distribution, respectively. An explicit representation of F is obtained by a function of independent random variables. We also derive an independent result on random partitions itself: a law of large numbers for the restricted uniform measure. This is a joint work with Ke Wang

April 15, 2015

Paths to precision medicine: Exploratory Subgroup Analysis in Clinical Trials

Xin Huang : 4 p.m. in SEO 636
Abstract Mechanistic relationships between the clinical outcome (efficacy or safety endpoints) versus putative biomarkers, clinical baseline and related predictors are usually unknown, and must be deduced empirically from experimental data. Such relationships enable the implementation of a personalized medicine strategy in clinical trials to help stratify patients in terms of disease progression, clinical response, treatment differentiation, etc. These relationships are often requires complex modelling to develop the prognostic and predictive signatures. For the purpose of easier interpretation and implementation in the clinical practice, defining a multivariate biomarker signature in terms of thresholds on the biomarker combinations is preferable. In this presentation, we propose some methods for developing such signatures in the context of continuous, binary and time-to-event endpoints. Results from simulations and case-study illustration will also be provided.

April 22, 2015

A Graphical Algorithm for the Nucleolus of Binary Assignment Games / Multivariate Final-Offer Arbitration

John Hardwick / Brian Powers : 4 p.m. in SEO 636
Abstract (Hardwick) An assignment game is a cooperative game with its player set divided into two groups, where a coalition receives a positive payoff if and only if it contains at least one player from both groups. In other words, it concerns a bipartite matching, with a payoff defined for each pair (payoffs for larger coalitions are determined additively). The 1994 paper by Solymosi and Raghavan gives an algorithm for finding the nucleolus (an optimal solution) of such a game, requiring O(n4) operations (assuming that n is the size of both groups). In this paper, we examine the special case in which the payoffs for each pair take only binary values, which we can think of as an indicator of whether or not the pair is compatible. For this case, we present a new algorithm which capitalizes on the graphical aspect of the previous one. This algorithm is shown to be an improvement, requiring only O(n3) operations. We also discuss sociological implications of this solution, considering each pair’s split of the payoff to be their relationship’s balance of power. (Powers) Final-Offer Arbitration is used in major league baseball to determine players' salaries, and has been used in various other wage disputes. In Final-Offer Arbitration, two contesting parties each provide a mediator with a final offer. The mediator must choose one of the two offers with no option for compromise. Under certain conditions, Nash equilibria are known for the single-variable case. Here we will explore solution points when players have multiple components to their offer, effectively bringing the dimension of the game to 2 or higher.

April 29, 2015

Panel discussion

Stat faculty and students : 4 p.m. in SEO 636
Abstract This panel discussion will give students an opportunity to have questions about various things (e.g., advantages and disadvantages of an academic career, general strategies for successful research, career opportunities outside of academia and how to prepare for them, etc) answered by faculty and senior graduate students.

Sept. 2, 2015

Organizational Meeting

Statistics Faculty & Students : 4 p.m. in SEO 636

Sept. 9, 2015

Scaling the Gibbs posterior

Nick Syring : 4 p.m. in SEO 636
Abstract In some applications, the relationship between the observable data and unknown parameters is described via a loss function rather than likelihood. In such cases, the standard Bayesian methodology cannot be used, but a Gibbs posterior distribution can be constructed by appropriately using the loss in place of a likelihood. Inference based on the Gibbs posterior is not straightforward, however, because the finite-sample performance is highly sensitive to the scale of the loss function. In this talk, I will propose a Gibbs Posterior Scaling (GPS) algorithm that adaptively selects the scaling in order to calibrate the corresponding Gibbs posterior credible regions. Two examples, namely, classification and quantile regression, are used to demonstrate that the Gibbs posterior with scale chosen by GPS produces valid interval estimates which are at least as efficient those obtained from other methods.

Sept. 16, 2015

Mock Oral Exam

Stat Faculty and Students : 4 p.m. in SEO 636

Sept. 23, 2015

Intermittency properties for a family of SPDEs driven by fractional noise

Daniel Conus : 4 p.m. in SEO 636
Abstract The talk will focus on results related to the notion of intermittency: i.e. the property that a random field develops large values ("high peaks") when time gets large. We will first present it with known examples and, then, describe how this phenomenon appears in the context of SPDEs. In particular, we will illustrate how the intermittent behavior of the solution to an SPDE depends on the type of driving noise in the case of stochastic heat and wave equations driven by fractional noise. In the latter case, the results are obtained via a Feynman-Kac representation of the moments similar to the one introduced in Dalang-Mueller-Tribe (2008). This is based on a joint work with Raluca Balan (Univ. of Ottawa).

Sept. 30, 2015

An ensemble distance measure of k-mer and Natural Vector for the phylogenetic analysis of multiple-segmented viruses

Hsin-Hsiung Huang : 4 p.m. in SEO 636
Abstract The Natural Vector combined with Hausdorff distance has been successfully applied for classifying and clustering multiple-segmented viruses. Additionally, k-mer methods also yield promising results for global genome comparison. It is not known whether combining these two approaches can lead to more accurate results. The author proposes a method of combining the Hausdorff distances of the 5-mer counting vectors and natural vectors which achieves the best classification without cutting off any sample. Using the proposed method to predict the taxonomic labels for the 2,363 NCBI reference viral genomes dataset, the accuracy rates are 96.95%, 94.37%, 99.41% and 93.82% for the Baltimore, family, subfamily, and genus labels, respectively. We further applied the proposed method to 48 isolates of the influenza A H7N9 viruses which have eight complete segments of nucleotide sequences. The single-linkage clustering trees and the statistical hypothesis testing results all indicate that the proposed ensemble distance measure can cluster viruses well using all of their segments of genome sequences.

Oct. 7, 2015

Design of Dose-response Clinical Trials

Naitee Ting : 4 p.m. in SEO 636
Abstract In the process of drug discovery and drug development, understanding the dose-response relationship is one of the most challenging tasks. It is also critical to identify the right range of doses in early stages of clinical development so that Phase III trials can be designed to confirm these doses. Usually at the beginning of Phase II, there is not a lot of available information to help guiding the study design. At this stage, Phase II clinical studies are needed to establish proof of concept (PoC), to identify a set of potentially effective and safe doses, and to estimate dose-response relationships. Challenges in designing these studies include: selection of the dose frequency and the dose range, choice of clinical endpoints or biomarkers, and use of control(s), among others. Consequences of bad Phase II study designs may lead to the delay of the entire clinical development program or the waste of R&D investment. Misleading results obtained from poor designs could cause a Phase III program to confirm a wrong set of doses, or to stop developing a potentially useful drug. Therefore, it is critical to consider an entire drug development plan, to make best use of all the available information, and to include all relevant experts in designing Phase II dose response clinical trials. This presentation discusses some of these considerations.

Oct. 14, 2015

Weighted Optimality Criteria and Design Search Algorithms

Jonathan Stallings : 4 p.m. in SEO 636
Abstract Standard design criteria like the A-, E-, and D-criterion implicitly assume the experimenter is equally interested in all estimable functions. Because of this, efficient designs under these criteria spread information evenly across the estimation space. In some cases, optimal designs can be analytically derived under these criteria but researchers are beginning to rely on design search algorithms to find these designs. These computer-generated designs are often found under the D-criterion because of its fast computations with point- and coordinate-exchange algorithms. However, the D-criterion is a poor assessment of a design when the goals of an experiment imply differential interest among the estimable functions. To reflect relative importance, Stallings and Morgan (Biometrika, 2015, in press) introduced general weighted optimality criteria, which assign weights to variances so that greater weight implies greater interest. These criteria are natural extensions of standard design criteria so that design search algorithms can be easily modified to perform optimization with respect to this new class of criteria. This talk first reviews the theory of general weighted optimality criteria and shows how the weighted analogues of standard criteria behave. A straightforward modification of typical design search algorithms is then shown to perform weighted optimization. The algorithm is implemented in SAS PROC OPTEX to find efficient blocked treatment-versus-control designs; unblocked and blocked factorial experiments that are focused on main effect estimation; and factorial experiments under a baseline parameterization.

Oct. 21, 2015

Time-fractional and L-Kuramoto-Sivashinsky (S)PDEs: two sides of the Brownian-time coin

Hassan Allouba : 4 p.m. in SEO 636
Abstract High order and fractional PDEs have become prominent in theory and in modeling many phenomena. We introduce two large classes of time-fractional and fourth order L-Kuramoto-Sivashinsky (L-KS) Stochastic PDEs. The L-KS PDE/SPDE class is connected to many pattern-formation phenomena. The latter class of time-fractional stochastic equations is related to noisy slow diffusion or diffusion in material with memory. We give comprehensive, sharp, and dimension-dependent Holder and modulus of continuity regularity results for both classes. One important theme of this talk—which is based on a series of our papers—is on the key role Brownian-time processes, their extensions, and their associated kernels play in giving a unifying explicit formulation and in capturing precise behaviors of these two important classes.

Oct. 28, 2015

Bayesian Graphical Models with Application to Integrate TCGA data

Yuan Ji : 4 p.m. in SEO 636
Abstract The Cancer Genomes Atlas (TCGA) data are unique in that multimodal measurements across genomics features, such as copy number, DNA methylation, and gene expression, are obtained on matched tumor samples. The multimodality provides an unprecedented opportunity to investigate the interplay of these features. Graphical models are powerful tools for this task that address the interaction of any two features in the presence of others, while traditional correlation- or regression-based models cannot. We introduce Zodiac, an online resource consisting of a large database containing nearly 200 million interaction networks of multiple genomics features produced by applying novel Bayesian graphical models on TCGA data through massively parallel computation. Setting a new way of integrating TCGA data, Zodiac, publically available at http://www.compgenome.org/ZODIAC, is expected to facilitate the generation of new knowledge and hypotheses by the community.

Nov. 4, 2015

Random Walks and Diffusions on Graphs

Melanie Pivarski : 4 p.m. in SEO 636
Abstract Random walks and diffusions are intimately connected through their relationship to the heat (diffusion) equation. We will look at their structures as well as results comparing large scale asymptotics for finitely generated groups and associated manifolds (Pittet & Saloff-Coste 2000) and local Dirichlet spaces (P. 2012).

Nov. 18, 2015

A measure of information in non-regular problems

Yi Lin : 4 p.m. in SEO 636
Abstract Fisher information plays an important role in statistics, from asymptotic efficiency, to optimal design of experiments, to the construction of default priors for Bayesian analysis. However, existence of Fisher information requires regularity conditions. What happens if the regularity conditions are not met? Is there an alternative measure of information that can be used in non-regular problems when the Fisher information does not exist? In this talk, I will present a generalization of the Fisher information to non-regular problems, based on the Hellinger distance, and discuss its properties and some examples. Hints about its application to optimal design of experiments may also be given.

Dec. 2, 2015

Spatial asymptotics for the parabolic Anderson models with generalized time-space Gaussian noise

Xia Chen : 4 p.m. in SEO 636
Abstract This work is concerned with the precise spatial asymptotic behavior for the parabolic Anderson equation $$\frac{\partial u}{\partial t}(t,x)=\frac{1}{2}\triangle u(t,x)+V(t,x)u(t,x),\quad\quad\mathrm{with}\ u(0,x)=u_0(x),$$ where the homogeneous generalized Gaussian noise $V(t,x)$ is, among other forms, white or fractional white in time and space. Associated with the Cole-Hopf solution to the KPZ equation, in particular, the precise asymptotic form $$\lim_{R\to+\infty}(\log R)^{-2/3}\log\max_{|x|\leq R}\, u(t,x)=\frac{3}{4}\sqrt[3]{\frac{2t}{3}}\quad \mathrm{a.s.}$$ is obtained for the parabolic Anderson model $\partial_t u=\frac{1}{2}\partial^2_{xx}u+\dot{W}u$ with the $(1+1)$-white noise $\dot{W}(t,x)$.

Jan. 20, 2016

Biomedical Big Data and Permanental Classification Approach

Jie Yang : 4 p.m. in SEO 636
Abstract The explosion in the availability of biomedical data is creating both great opportunities and challenges for collaborative research among clinicians, genomics and proteomics scientists, molecular biologists, and statisticians. On one hand, electronic medical records and genomic data of a large cohort of individuals are assembled and become available for health study researches. On the other hand, the combined data are extremely high-dimensional and also becoming bigger and bigger, especially the genomic part. As one of the most critical application areas with the biomedical big data, precision medicine refers to precisely classifying individuals into subpopulations according to their susceptibility to a particular disease and precisely tailoring of medical treatments to subcategories of the disease. Achieving the goals of precision medicine requires combining data across multiple formats and developing novel, sophisticated statistical methods. Our permanental classification approach recently developed is capable of handling high-dimensional classification problems. It provides a promising solution for biomedical high-dimensional data, implemented using the most popular open source statistical software, R.

Jan. 27, 2016

Optimal Bayesian posterior concentration rates with empirical priors

Ryan Martin : 4 p.m. in SEO 636
Abstract A Bayesian approach provides a technically straightforward procedure to produce inference on high- and even infinite-dimensional parameters in complex models. Of course, the choice of a prior is always an issue and, especially in high-dimensional problems, the prior has a non-trivial effect. One attempt use data to help select an appropriate prior is <i>empirical Bayes</i> but, unfortunately, this approach does not lead to any theoretical guarantees that the posterior will behave properly. In this talk I will introduce a very simple strategy that incorporates data into the prior in such a way that the corresponding posterior distribution has optimal, even adaptive, concentration rates. Some illustrations of the general theory will also be presented. (This is joint work with Stephen Walker at University of Texas--Austin.)

Feb. 3, 2016

Einstein relation and steady states for the random conductance model

Xiaoqin Guo : 4 p.m. in SEO 636
Abstract The Einstein relation describes the relation between the response of a system to a perturbation and its diffusivity at equilibrium. It states that the derivative (with respect to the strength of the perturbation) of the velocity equals the diffusivity. In this talk we consider random walks in iid random conductances on the integer lattice $Z^d$. We show that when $d\ge 3$, the invariant measure for the environment viewed from the particle has a first order expansion in terms of the perturbation. The Einstein relation will follow as a corollary of this expansion. This talk is based on a joint work with N. Gantert and J. Nagel.

Feb. 17, 2016

Statistical Inference for Stochastic PDEs

Igor Cialenco : 4 p.m. in SEO 636
Abstract We consider a parameter estimation problem for finding the drift coefficient for a large class of parabolic Stochastic PDEs driven by additive or multiplicative noise. In the first part of the talk, we derive several different classes of estimators based on the first N Fourier modes of a sample path observed continuously on a finite time interval. In the second part of the talk we will investigate the simple hypothesis testing problem for the drift coefficient for stochastic fractional heat equation driven by additive noise. We introduce the notion of asymptotically the most powerful test, and find explicit forms of such tests in two asymptotic regimes: large time asymptotics, and increasing number of Fourier modes. Also, we will discuss how to estimate and control the Type~I and Type~II errors. Finally, we illustrate the theoretical results by some numerical examples/simulations.

Feb. 24, 2016

Statisticians come in from the cold: BigQuery, R, and Shiny

Troy Hernandez : 4 p.m. in SEO 636
Abstract There's been a lot of hand-wringing over the fumbling of the statistics field in harnessing the popularity of data science and the "big data revolution". This was most famously addressed in the editorial from the previous ASA president Marie Davidian, ["Aren't We Data Science?"](http://magazine.amstat.org/blog/2013/07/01/datascience/). While I have to continue to inform peers and colleagues that R is indeed a Turing complete programming language (and not just a statistical computing environment), the data science hype has created technologies that enable any statistician get into the data science game; most notably Shiny, and cloud-based databases. My talk will focus on weaving these tools together in a way that allows statisticians to say, ["All your data science are belong to us."](http://knowyourmeme.com/memes/all-your-base-are-belong-to-us) Additionally, I will provide an update on the issue of diversity in the tech industry from the frontline.

March 2, 2016

Modeling between- and within-subject variances using mixed-effects location scale models for intensive longitudinal data

Donald Hedeker : 4 p.m. in SEO 636
Abstract Intensive longitudinal data are increasingly encountered in many research areas. For example, ecological momentary assessment and/or experience sampling methods are often used to study subjective experiences within changing environmental contexts. In these studies, up to 30 or 40 observations are usually obtained for each subject over a period of a week or so. Because there are so many measurements per subject, one can characterize a subject's mean and variance and can specify models for both. In this presentation, we focus on an adolescent smoking study using ecological momentary assessment where interest is on characterizing changes in mood variation. We describe how covariates can influence the mood variances and also extend the statistical model by adding a subject-level random effect to the within-subject variance specification. This permits subjects to have influence on the mean, or location, and variability, or (square of the) scale, of their mood responses. These mixed-effects location scale models have useful applications in many research areas where interest centers on the joint modeling of the mean and variance structure.

March 9, 2016

Noisy differential equations with power type coefficients

Samy Tindel : 4 p.m. in SEO 636
Abstract We are interested in this talk in ordinary differential equations with a noisy term and a diffusion type coefficient of the form $|x|^{a}$, with a constant a smaller than 1. This kind of equation has a long story in stochastic analysis. We will first review some of the efforts made by Yamada and Watanabe in this direction (when the equation is driven by a Brownian motion), as well as more some recent developments concerning stochastic PDEs. We will then introduce two extensions of Young's integral which allows to handle the case of equations driven by a Gaussian signal whose paths are Hölder continuous with Hölder exponent greater than 1/2. Notice that only existence results are obtained, the uniqueness part being still widely open. This presentation is based on a joint work with J. León and D. Nualart.

March 16, 2016

TBA

Rich Bu : 4 p.m. in SEO 636
Abstract There are two topics in this talk. 1. Applications of statistical models in the consumer lending industry 2. The job market in a nutshell in consumer lending modeling

March 30, 2016

Symmetric Random Walks on Tetrahedra and Octahedra

Jyotirmoy Sarkar : 4 p.m. in SEO 636
Abstract We consider a symmetric random walk on the vertices of a tetrahedron or an octahedron. Starting from the origin, at each step the random walk moves to one of the vertices adjacent to the current vertex with equal probability. We find the distribution, or at least the mean and the variance, of the number of steps needed to (1) return to origin, (2) visit all vertices, and (3) return to origin after visiting all vertices. We also obtain the distributions of (i) the number of vertices visited before return to origin, (ii) the last vertex visited, and (iii) the number of vertices visited while returning to origin after visiting all vertices.

April 6, 2016

An Affine-Invariant Bayesian Cluster Process with Split-Merge Gibbs Sampler

Hsin-Hsiung Huang : 4 p.m. in SEO 636
Abstract We develop a clustering algorithm which does not requires knowing the number of clusters in advance. Furthermore, our clustering method is rotation-, scale- and translation-invariant. We call it ``Affine-invariant Bayesian (AIB) process". A highly efficient split-merge Gibbs sampling algorithm is proposed. Using the Ewens sampling distribution as prior of the partition and the profile residual likelihoods of the responses under three different covariance matrix structures, we obtain inferences in the form of a posterior distribution on partitions. The proposed split-merge MCMC algorithm successfully and efficiently estimate the partition. Our experimental results indicate that the AIB process outperforms other competing methods. In addition, the proposed algorithm is irreducible and aperiodic, so that the estimate is guaranteed to converge to the true partition.

April 20, 2016

A flexible Bayesian nonparametric model for predicting future insurance claims

Liang Hong : 4 p.m. in SEO 636
Abstract Accurate prediction of future claims is a fundamentally important problem in insurance. The Bayesian approach is natural in this context, as it provides a complete predictive distribution for future claims. The classical credibility theory provides a simple approximation to the mean of that predictive distribution as a point-predictor, but this approach ignores other features of the predictive distribution, such as spread, that would be useful for decision-making. Unfortunately, these other features are more sensitive to the choice of loss model and prior distribution, so a flexible nonparametric Bayesian model is desirable. In this paper, we propose a Dirichlet process mixture of log-normals model and discuss the theoretical properties and computation of the corresponding predictive distribution. Numerical examples demonstrate the benefit of our model compared to some existing insurance loss models, and an R code implementation of the proposed method is also provided.

April 27, 2016

Panel Discussion

Stat faculty and students : 4 p.m. in SEO 636
Abstract This panel discussion will give students an opportunity to have questions about various things (e.g., advantages and disadvantages of an academic career, general strategies for successful research, career opportunities outside of academia and how to prepare for them, etc) answered by faculty and senior graduate students.

Aug. 31, 2016

Organizational meeting

Stat faculty and graduate students : 4 p.m. in SEO 636

Sept. 7, 2016

Competing Brownian particles

Andrey Sarantsev : 4 p.m. in SEO 636
Abstract Consider a finite or infinite system of Brownian particles on the real line. Each particle moves as a Brownian motion with drift and diffusion coefficients depending on its current rank relative to other particles. These systems were introduced in Banner, Fernholz, Karatzas (2005). Since then, extensive theory was developed for finite systems. However, infinite systems proved to be much more difficult. We survey the latest results.

Sept. 14, 2016

Coverage Probability and Exact Inference

Bikas Sinha : 4 p.m. in SEO 636
Abstract Abstract : With reference to 'point estimation' of a real-valued parameter $\theta$ involved in the distribution of a real-valued random variable $X$, we consider a sample size $n$ and an underlying exact-sense unbiased estimator ${\hat{\theta}}_n$ of $\theta$ for every $n = k, k+1, k+2, ...\ldots$ where $k$ is the minimum sample size for existence of an exact-sense unbiased estimator of $\theta$. We wish to investigate exact small sample properties of the sequence of estimators considered here. This we study by considering what is termed as 'Coverage Probability (CP)' and defined as $CP(n, c)=P[-c < {\hat{\theta}}_n - \theta < c]$. It is desired that the sequence $[CP(n, c); n=k, k+1, k+2, ...\ldots]$ behaves like an increasing sequence for every $c>0$. We may note that we are asking for a property beyond 'consistency' of a sequence of estimators. In this presentation we will discuss several interesting features of the behavior of the $CP(n, c)$.

Sept. 21, 2016

Credit Risk Management of Non-secured Lending

Yan Chang : 4 p.m. in SEO 636
Abstract I will go over some of principles and typical approaches of credit risk management in a financial company.

Sept. 28, 2016

Some Recent Developments on the Applications of Evolutionary Algorithm in the Statistical Optimization

Frederick Phoa : 4 p.m. in SEO 636
Abstract Nature-inspired metaheuristic methods, like the particle swarm optimization and many others, enjoys fast convergence towards optimal solution via a series of inter- particle communication. Such methods are common for the optimization problem in engineering, but few in statistics problem. It is especially difficult to implement in some fields of statistics as the search spaces are mostly discrete, while most natural heuristic methods require continuous search domains. This talk introduces a new method called the Swarm Intelligence Based (SIB) method for optimization in statistics problems, featuring the searches within discrete space. Such fields include experimental designs, community detection, change-point analysis, variable selection, etc. The SIB method is a nature-inspired metaheuristic method that includes several operations. This method is advantageous over the traditional particle swarm optimization and many other heuristic approaches in the sense that it is ready for the search of both continuous and discrete domains, and its global best particle is guaranteed to monotonically move towards the optimum. The SIB method is demonstrated in several examples. Several extensions from the standard framework are also discussed at the end of this talk.

Oct. 5, 2016

New Approaches to Fast Approximate Bayesian Nonparametric Inference

George Karabatsos : 4 p.m. in SEO 636
Abstract Dirichlet process (DP) mixture models, as well as models with mixture distribution assigned a general Bayesian nonparametric (BNP) prior distribution on the space of probability measures, are widely-applied and flexible models that can provide reliable statistical inferences complex data. For such Bayesian mixture models, in practice, posterior inferences are usually conducted using MCMC, which however, is prohibitively slow for large data sets. Also for such models, prior specification can be non-trivial in practice. As alternatives to MCMC, I consider two new approaches to fast and approximate BNP inference for large data sets. First, I show that if the ordinary least-squares (OLS) estimator of the linear regression coefficients is specified as a functional of the DP posterior distribution, then this functional has posterior mean given by an observation-weighted ridge regression estimator, with ridge (coefficient shrinkage) parameter given by the DP precision parameter; and has a heteroscedastic-consistent posterior covariance matrix. This result is based on the multivariate delta method applied to prior-informed bootstrap distribution approximation to the DP posterior. Second, I consider an approximation to the BNP (infinite) mixture model that I introduced and studied in several articles, defined by ordinal regression mixture weights.The approximate model is defined by a (large) finite mixture, with each component distribution multiplied by a histogram bin indicator function. I show that posterior inference with this approximate BNP model can be conducted by iteratively-reweighted least squares estimation for the mixture weight parameters, and least-squares estimation for the component densities, all involving computations that are orders of magnitude faster that MCMC-based inference of the original mixture model. This is also true for a version of the approximate model that is defined by an ordinal regression of DPs. I illustrate the two approximate BNP methods through the analysis of real data sets.

Oct. 12, 2016

Small-time asymptotics of subRiemannian Hermite functions

Tai Melcher : 4 p.m. in SEO 636
Abstract As in the Riemannian setting, a subRiemannian heat kernel is controlled by the geometry of the underlying manifold. In particular, the asymptotic behavior of the kernel can reveal certain geometric and topological data. We study the logarithmic derivatives of subRiemannian heat kernels in some cases and show that, under appropriate scaling, they converge to their analogues on stratified groups. This gives one quantification of the now standard idea that stratified groups play the role of the tangent space to subRiemannian manifolds. This is joint work with Joshua Campbell.

Oct. 19, 2016

Model-based approaches to learn partitions from data

Dongxiao Zhu : 3 p.m. in SEO 636
Abstract In multi-class classification, different classes may relate to different feature groups. In this talk, I will present a class-conditional regularization of the multinomial logistic model to enable the discovery of class-specific feature groups. I will also present an efficient cyclic block coordinate descent based algorithm to solve the model. In another work, I will introduce a novel joint mixture model framework to estimate cluster size distribution, particularly for over-dispersed (high variance) ones, together with cluster compactness (density). Our methods are sufficiently flexible and general to be applied to multiple application domains, such as social networks, image segmentation, natural language processing and bioinformatics.

Recent advances in optimal design for correlated Processes

Milan Stehlik : 4 p.m. in SEO 636
Abstract Since 2004 there were many results obtained regarding the determination of optimal designs for models with correlated errors. This task is substantially more difficult than in case of iid errors and for this reason not so well developed. Stochastic process with parametrized mean and covariance is observed over a compact set. The information obtained from observations is measured through the information functional (defined on the Fisher information matrix). The role of equidistant designs has been recognized; e.g. such designs have been proved to be optimal for parameter of trend of stationary Ornstein-Uhlenbeck process, also for nonstationary Ornstein-Uhlenbeck process, both for prediction and estimation. We can conclude that if only trend parameters are of interest, the designs covering more-less uniformly the whole design space are rather efficient when correlation decreases exponentially. This concept is also valid for so called monotonic set designs. We will concentrate on several important issues regarding regularity conditions for quality of ``plug-in" approach from iid case. Namely, 1) relaxing the continuity of covariance. We will introduce the regularity conditions for isotropic processes with semicontinuous covariance such that increasing domain asymptotic is still feasible, however more flexible behavior may occur here. In particular, the role of the nugget effect will be illustrated. 2) regarding quality of approximation of inverse information matrix by variance-covariance, Pazman (2007) discusses theoretical background and formulated important conditions, which are fundamental. Zhu and Stein (2005) made simulations experiments. Finally, application in troposphere methane modelling will be illustrating the developed methods.

Oct. 26, 2016

Design and Analysis of Clinical Studies using Restricted Mean Survival Time

Lihui Zhao : 4 p.m. in SEO 636
Abstract For a study with an event time as the endpoint, its survival function contains all the information regarding the temporal, stochastic profile of this outcome variable. The survival probability at a specific time point, say t, however, does not transparently capture the temporal profile of this endpoint up to t. An alternative is to use the restricted mean survival time (RMST) at time t to summarize the profile. The RMST is the mean survival time of all subjects in the study population followed up to t, and is simply the area under the survival curve up to t. The advantages of using such a quantification over the survival rate have been discussed in the setting of a fixed-time analysis. In this research, we generalize this approach by considering a curve based on the RMST over time as an alternative summary to the survival function. Inference, for instance, based on simultaneous confidence bands for a single RMST curve and also the difference between two RMST curves are proposed. The latter is informative for evaluating two groups under an equivalence or noninferiority setting, and quantifies the difference of two groups in a time scale. In addition, we extend RMET to the setting of multiple endpoints, which includes classical competing risks and semi-competing risks. The methods are illustrated with the data from two clinical trials.

Nov. 2, 2016

Overlaps and Pathwise Localization in the Anderson Polymer Model

Mike Cranston : 4 p.m. in SEO 636
Abstract We consider large time behavior of typical paths under the Anderson polymer measure. If $P^x_\kappa$ is the measure induced by rate $\kappa,$ simple, symmetric random walk on $\mathbb{Z}^d$ started at $x,$ this measure is defined as \[d\mu^x_{\kappa,\beta,T}T(X)={Z_{\kappa,\beta,T}}^{-1} \exp\left\{\beta\int_0^T dW_{X(s)}(s)\right\}dP^x_\kappa(X)\] where $\{W_x:x\in \mathbb{Z}^d\}$ is a field of $iid$ standard, one-dimensional Brownian motions, $\beta>0, \kappa>0$ and $Z_{\kappa,\beta,t}(x)$ the normalizing constant. We establish that the polymer measure gives a macroscopic mass to a small neighborhood of a typical path as $T \to \infty$, for parameter values outside the perturbative regime of the random walk, giving a pathwise approach to polymer localization, in contrast with existing results. The localization becomes complete as $\frac{\beta^2}{\kappa}\to\infty$ in the sense that the mass grows to 1. The proof makes use of the overlap between two independent samples drawn under the Gibbs measure $\mu^x_{\kappa,\beta,T}$, which can be estimated by the integration by parts formula for the Gaussian environment. Conditioning this measure on the number of jumps, we obtain a canonical measure which already shows scaling properties, thermodynamic limits, and decoupling of the parameters. This talk is based on joint work with Francis Comets.

Nov. 9, 2016

Bayesian Variable Selection in Complex Linear and Lifetime Models

Sanjib Basu : 4 p.m. in SEO 636
Abstract We consider the question of variable selection in complex models. This is often a difficult problem due to the inherent nonlinearity of the models and the resulting non-conjugacy in their Bayesian analysis. Bayesian variable selection in lifetime data models often utilize cross-validated predictive model selection criteria which can be relatively easy to estimate for a given model. However, the performances of these criteria are not well-studied in large-scale variable selection problems and, evaluation of these criteria for each model under consideration can be difficult to infeasible. An alternative criterion is based on the highest posterior model but its implementation is difficult in non-conjugate lifetime models. In this presentation, we compare the performances of these different criteria in complex lifetime data models including models with limited failure. We also propose an efficient variable selection method and illustrate its performance in simulation studies and real example

Nov. 16, 2016

Spatial SIR model and superprocesses

Si Tang : 4 p.m. in SEO 636
Abstract The classic Susceptible-Infected-Resistant (SIR) model, due to Kermack & McKendrick (1927), describes the spread of an infectious disease in an infinite, homogeneous population using a system of ordinary differential equations. In this talk, I will focus on stochastic SIR models, where the population is of size $N$ and transmissions only occurs locally. In these models, the sizes of infected clusters are characterized by the excursion lengths of a continuous stochastic process, denoted by $W_t$. In particular, in the mean-field SIR case, $W_t$ is a reflected (at 0) Brownian motion with negative drift; in the spatial case, $W_t$ is a reflected Brownian motion trimmed by a Poisson point process whose intensity is determined by a super Brownian motion with time-and-location dependent killing.

Nov. 30, 2016

Advanced Marketing Analytics: Data Science at Precima

Ella Revzin : 4 p.m. in SEO 1227
Abstract This talk will introduce students to Precima and to Marketing Data Science. Precima is a Marketing Analytics company that specializes in data driven products and services that help retailers and manufacturers drive sales growth and boost profitability. I will provide an overview of the company and its Data Science practice, with a focus on Targeted Marketing. For the latter, I’ll review a typical data acquisition and modeling process. Finally, I’ll describe a typical week in the life of a Precima Data Scientist and give information on what we look for when we hire for internships and full time positions.

Feb. 15, 2017

Deterministic Methods for Stochastic Dynamics

Jinqiao Duan : 4 p.m. in SEO 636
Abstract Dynamical systems arising in engineering and science are often subject to random fluctuations. The noisy fluctuations may be Gaussian or non-Gaussian, which are modeled by Brownian motion or α-stable Levy motion, respectively. Non-Gaussianity of the noise manifests as nonlocality at a “macroscopic” level. Stochastic dynamical systems with non-Gaussian noise (modeled by α-stable Levy motion) have attracted a lot of attention recently. The non-Gaussianity index α is a significant indicator for various dynamical behaviors. The speaker will overview recent advances in non-Gaussian stochastic dynamical systems, highlighting deterministic and numerical methods, including analysis and simulation of mean exit time and escape probability. Some materials are taken from the speaker's new book “An Introduction to Stochastic Dynamics” (Cambridge University Press, 2015) .

Feb. 22, 2017

Functional Coefficient Time Series Models with Trending Regressors

Tingting Cheng : 4 p.m. in SEO 636
Abstract We study a functional coefficient time series model with trending regressors, where the coefficients are unknown functions of time and random variables. We propose a local linear estimation method to estimate the unknown coefficient functions. An asymptotic distribution of the proposed local linear estimator is established under mild conditions. A test procedure is developed to test the null hypothesis that the functional coefficients take particular parametric forms. For practical use, we further propose a Bayesian approach to select bandwidths involved in this local linear estimator. Several numerical examples are provided to examine the finite sample performance of the proposed local linear estimator and the test procedure. The results show that the local linear estimator works well and the proposed test has satisfactory size and power. In addition, simulation studies show that the Bayesian bandwidth selection method is better than cross–validation method. Furthermore, we employ the functional coefficient model to study the relationship between consumption per capita and income per capita in U.S. and the results show that functional coefficient model with our proposed local linear estimator and Bayesian bandwidth selection method performs best in both in–sample fitting and out–of–sample forecasting.

March 29, 2017

Some statistical considerations in High-Throughput-Screening data evaluations in drug discovery

Dr. Viswanath Devanarayan : 4 p.m. in SEO 636
Abstract In High-Throughput-Screening efforts during the drug discovery process, hundreds of thousands of compounds are tested to identify promising drug candidates that modulate specific gene targets. These drug candidates may ultimately serve as therapeutic candidates for some disease indications of interest. Critical decisions related to compound selection and prioritization are made based on fairly limited data, and therefore rely greatly on data quality and reproducibility. Standard statistical metrics and methods in textbooks do not directly apply for these evaluations. This presentation will provide an overview of some statistical measures that were developed specifically for this application. The content of this presentation will be very practical and data-driven, and hence will be suitable for a broad audience.

April 5, 2017

Propagation of critical behavior for unitary invariant plus GUE random matrices

Karl Liechty : 4 p.m. in SEO 636
Abstract It is a well known and celebrated fact that the eigenvalues of random Hermitian matrices from a unitary invariant ensemble form a determinantal point process with correlation kernel given in terms of a system of orthogonal polynomials on the real line. It is a much more recent result that the eigenvalues of the sum of such a random matrix with a matrix from the Gaussian unitary ensemble (GUE) also forms a determinantal point process, with the kernel given in terms of the Weierstrass transform of the original kernel. I'll talk about the case in which the limiting distribution of eigenvalues is critical in the sense that there is a non-generic scaling limit for the correlation kernel, and discuss the effect of a Gaussian perturbation on the limiting critical kernel. This is joint work with Tom Claeys, Arno Kuijlaars, and Dong Wang.

April 7, 2017

Crossover Designs and Related Open Research Problems

Wei Zheng : 11 a.m. in TH300

April 12, 2017

Excursion landscape

Ju-Yi Yen : 4 p.m. in SEO 636
Abstract In this talk, we study the process obtained from a Brownian bridge after excising all the excursions below the waterline level which reach zero. Three variables of interest are the maximum of this process, the value where this maximum is attained, and the total length of the excursions which are excised. Our analysis relies on some interesting transformations connecting Brownian path fragments and the 3-dimensional Bessel process.

April 19, 2017

The exponential transform of a path

Xi Geng : 4 p.m. in SEO 636
Abstract The exponential transform of a vector-valued path, also known as the signature of a path, is the formal sequence of associated iterated path integrals. It is widely believed (and surprisingly) that the signature contains essentially all information about the underlying path. In this talk, we will prove that every (rough) path is uniquely determined by its signature up to certain tree-like equivalence. Moreover, looking into its probabilistic counterpart, we will obtain stronger uniqueness results for sample paths of Gaussian processes by applying the technique of Malliavin's calculus. This part inspires the development of a universal way to reconstruct every rough path from its signature.

April 26, 2017

P-SVM: Efficient Parameter Selection for Support Vector Machines with Gaussian Kernels

Prof. Hsin-Hsiung Huang : 4 p.m. in SEO 636
Abstract Support Vector Machines (SVM) classifier is a popular classification method. However, most users may not well take tuning parameters selection because this step is time consuming. In practice, the tuning parameters are chosen by evaluating parameter candidates via cross validation. It is shown that the performance of SVM is sensitive to the values of tuning parameters. In some cases, SVM performs poorly due to the values of tuning parameters. However, selection of parameter values for SVM often relies on inefficient approaches such as extensive cross validation. To get around the problem, users may resort to anecdotal methods or default values set by software developers. However, these methods may compromise performance of classification accuracy. In this research, we propose an efficient algorithm called P-SVM for selecting the parameter pair, (gamma,C), of SVM with Gaussian kernels on metric data. P-SVM searches only a handful of percentiles of the squared Euclidean distances of data points to select the best pair of parameter values. Our motivation case study of business intelligence categorization demonstrates that P-SVM achieved a signi cant improvement in precision, recall, F-measure, and AUC from the default parameter values settled in Weka, a widely used data mining software. Applications of both simulation and publicly-available datasets also demonstrate that P-SVM achieves substantial improvement in computational time without loss of much classification accuracy.

May 3, 2017

Recent advances in crossover designs and related studies

Prof. Wei Zheng : 4 p.m. in SEO 636
Abstract Crossover design is a design of experiments, where a subject receives a sequence of various treatment over a period of time points. While it provides the within subject comparison between treatment effects, the potential carryover effect in the model makes the study of optimal crossover designs quite complicated. Such study was initiated by Hedayat and Afsarinejad (1978), and many researchers have contributed to the general theory for optimal designs. Among them, Kushner (1997) developed very elegant results for the optimality conditions in the approximate design theory. My talk will mainly focus on this approach and talk about some recent progress as well as future challenges. I will also share some of my own thoughts of how to tackle these problems.

Aug. 7, 2017

Semi-parametric method for non-ignorable missing in longitudinal data using refreshment samples

Dr. Lan Xue : 3 p.m. in SEO 636
Abstract Missing data is one of the major methodological problems in longitudinal studies. It not only reduces the sample size, but also can result in biased estimation and inference. It is crucial to correctly understand the missing mechanism and appropriately incorporate it into the estimation and inference procedures. Traditional methods, such as the complete case analysis and imputation methods, are designed to deal with missing data under unverifiable assumptions of MCAR and MAR. The purpose of this talk is to identify and estimate missing mechanism parameters under the non-ignorable missing assumption utilizing the refreshment sample. In particular, we propose a semi-parametric method to estimate the missing mechanism parameters by comparing the marginal density estimator using Hirano?s two constraints (Hirano et al. 1998) along with additional information from the refreshment sample. Asymptotic properties of semi-parametric estimators are developed. Inference based on bootstrapping is proposed and verified through simulations.

Aug. 30, 2017

Intrinsic Ultracontractivity of Laplacian and Fractional Laplacian Perturbed by Non-local Operator

Yinghui Shi : 4 p.m. in SEO 636
Abstract The intrinsic ultracontractivity (Abbr. I.U.) of the subprocesses $X_D^b$ of two kinds of special Markov processes $X^b$ upon leaving any bounded open set $D \subset \mathbb{R}^d$ will be given in this talk. Here the processes $X^b$ are associated with the operator $\mathcal{L}^b= \Delta^{\alpha/2} +\mathcal{S}^b$ with $d \geq 1$ and $0 < \beta < \alpha \leq 2$, where $$\mathcal{S}^bf(x):=\int_{\mathbb{R}^d}(f(x+z)-f(x)-\nabla f(x)\cdot z \mathbb{1}_{\{|z|\leq 1\}})\frac{b(x,z)}{|z|^{d+\beta}}dz$$ and $b(x,z)$ is a bounded Borel function on $\mathbb{R}^d\times \mathbb{R}^d$ with $b(x,z)=b(x,-z)$ for $x,z\in \mathbb{R}^d$. The operator $\mathcal{L}^b$ can be seen as the Laplacian ($\alpha = 2$) or the fractional Laplace operator ($0<\alpha <2$) with a lower order perturbation $\mathcal{S}^b$. Our main results are proved under the frame of the I.U. for non-symmetric Levy processes. We discuss the transition density function for $X_D^b$ firstly and then we get its dual process under some reference measure. At last, we prove that I.U. stands under the conditions as follows: for any compact subset $K, L \subset\mathbb{R}^d$, $\inf_{x\in K}\inf_{z\in L}b(x,z)>0$ in the case of $\alpha=2$ and $\inf_{x\in K}\inf_{z\in L}(1+\frac{b(x,z)}{\mathcal{A}(d,\alpha)})|z|^{\alpha-\beta}>0$ in the case of $0<\alpha<2$. (Joint work with Yi, Bingji and Song, Renming)

Sept. 6, 2017

Organizational meeting

Stat faculty and graduate students : 4 p.m. in SEO 636
Abstract TBA

Sept. 13, 2017

Inventory Pooling under Multivariate Fat-Tail Demands

Prof. Zhen Liu : 4 p.m. in SEO 636
Abstract We study the classic inventory pooling problem by Eppen (1979) under a special class of multivariate fat-tail distribution: Normal Inverse Gaussian (NIG) demands to better fit real-world demand data. We obtain the optimal inventory level in a closed form by employing standardized NIG density function, and express the optimal expected costs in terms of unit NIG loss function. In addition to independent and identically distributed demands, our results complement Bimpikis and Markakis (2015) by considering correlated demands. We further discuss the transshipment problem of Dong and Rudi (2004) under NIG demands.

Sept. 27, 2017

Weighted limit theorems and applications

Yanghui Liu : 4 p.m. in SEO 636
Abstract The term “limit theorem” is associated with a multitude of statements having to do with the convergence of probability distributions of sums of increasing number of random variables. Given that a limit theorem result holds, “weighted limit theorem” considers the asymptotic behavior of the corresponding weighted sums. The weighted limit theorem problem has drawn a lot of attention in recent articles due to its key role in topics such as parameter estimations, Ito’s formula in law, time-discrete numerical schemes, and normal approximations, and various “unexpected” weighted limit theorems have been discovered since then. The purpose of this talk is to introduce a general framework and a transferring principle for this problem, and to provide improvement of the existing results in a few aspects.

Oct. 4, 2017

Joint Estimation of Fractal Indices for Bivariate Gaussian Processes

Yimin Xiao : 4 p.m. in SEO 636
Abstract Multivariate (or vector-valued) stochastic processes are important in probability, statistics and various scientific areas as stochastic models. In recent years, there has been increasing interest in investigating their statistical inference and prediction. In this talk, we study the problem for estimating jointly the fractal indices of a bivariate Gaussian process. These indices not only determine the smoothness of each component process, fractal behavior of the whole process, but also play important roles in characterizing the dependence structure among the components. Under the infill asymptotics framework, we establish joint asymptotic results for the increment-based estimators for bivariate fractal indices. Our main results show the effect of the cross dependence structure on the performance of the estimators. This is a joint paper with Yuzhen Zhou.

Oct. 18, 2017

Data Science 2.0

Dan Spillane : 4 p.m. in SEO 636
Abstract What's the purpose of Data Science anyway? In this discussion we'll explore how we need to turn data science upside-down to create the real value of this powerful trade. We need to push desired (business, social, economic...) outcomes to the forefront (the hypothesis) and leverage data, data platforms and AI to develop the questions we don't even know to ask and then help answer. We need to be data pioneers not just data engineers. Looking forward to a fruitful and living dialogue on Data Science 2.0.

Oct. 25, 2017

Simultaneous confidence bands in time-varying coefficient models

Sayar Karmakar : 4 p.m. in SEO 636
Abstract The term "time-varying(tv) coefficient model" refers to the framework of time series and regression models where the unknown coefficients vary across time. Deviation from the constancy of parameters is more natural due to the effect of several external factors/events or sometimes simply for a very long time-horizon. For the past two decades, tv regression models received considerable attentions however not much was done for  conditional heteroscedastic(CH) models until very recently.   The purpose of this talk is to introduce an unanimous framework to combine the treatments for tv linear regression, tv generalized regression and tv time-series models. Local linear M-estimation is used to estimate the unknown curves. We obtain a Bahadur representation of these estimated curves and use it to find the simultaneous confidence bands. To circumvent the logarithmic convergence rate of the theoretical bands, a Bootstrap method is proposed using an optimal Gaussian approximation. Some simulations for  tvARCH and tvGARCH models and analysis of some stock market datasets are presented. This is a joint work with Stefan Richter and Wei Biao Wu.

Nov. 1, 2017

Statistics in Marketing Research

Yunxiao He : 4 p.m. in SEO 636
Abstract The marketing industry produces and consumes an enormous amount of data. This talk will provide a quick overview of the marketing industry with focus on illustrating how data and statistical tools can be leveraged in connecting brands and consumers.

Nov. 22, 2017

The tail asymptotics of the Brownian signature

Xi Geng : 3 p.m. in SEO 636
Abstract In the groundbreaking work of B. Hambly and T. Lyons (Uniqueness for the signature of a path of bounded variation and the reduced path group, Ann. of Math., 2010), it has been conjectured that the geometry of a tree-reduced bounded variation path can be recovered from the tail asymptotics of its associated sequence of iterated path integrals. While this conjecture is still remaining open in the general deterministic case, in this talk we investigate a similar problem in the probabilistic setting for Brownian motion. It turns out that a martingale approach applied to the hyperbolic development of Brownian motion allows us to extract useful information from the tail asymptotics of Brownian iterated integrals, which can be used to determined the Brownian rough path along with its natural parametrization uniquely. This in particular strengthens the existing uniqueness results in the literature.

Causality in the joint analysis of longitudinal and survival data

Lei Liu : 4 p.m. in SEO 636
Abstract In many biomedical studies, disease progress is monitored by a biomarker over time, e.g., repeated measures of CD4, hemoglobin level in end stage renal disease (ESRD) patients. The endpoint of interest, e.g., death or diagnosis of a specific disease, is correlated with the longitudinal biomarker. The causal relation between the longitudinal and time to event data is of interest. In this paper we examine the causality in the analysis of longitudinal and survival data. We consider four questions: (1) whether the longitudinal biomarker is a mediator between treatment and survival outcome; (2) whether the biomarker is a surrogate marker; (3) whether the relation between biomarker and survival outcome is purely due to an unknown confounder; (4) whether there is a mediator moderator for treatment. We illustrate our methods by data from two clinical trials: an AIDS study and a liver cirrhosis study.

Dec. 6, 2017

Statistical methods for compositional data analysis with application in metagenomics

Hongmei Jiang : 4:15 p.m. in SEO 636
Abstract Metagenomics is a powerful tool to study the microbial organisms living in various environments. The abundance of a microorganism or a taxon is usually estimated using relative proportion or percentage in sequencing-based metagenomics studies. Due to the constraint of the sum of the relative abundances being 1 or 100%, standard conventional statistical methods may not be suitable for metagenomics data analysis. In this talk we will discuss characterization of the association between microbiome and disease status and variable selection in regression analysis with compositional covariates. Current statistical and computational methods that are being developed to analyze the metagnoimcs data and the challenges will also be highlighted.

Feb. 7, 2018

Bayesian Experimental Design and Hierarchical Model for Quantitative and Qualitative Responses

Lulu Kang : 4 p.m. in SEO 636
Abstract In many science and engineering systems both quantitative and qualitative output observations are collected. For short, we call such a system QQ system. In this talk, I will talk about a systematical approach for the experimental design and data analysis for the QQ system. Classic experimental design methods are not suitable here because they often focus on one type of responses. We develop both Bayesian D and A-optimal design methods for experiments with one continuous and one binary responses. Both noninformative and conjugate informative prior distributions on the unknown parameters are considered. The proposed design criterions has meaningful interpretations in terms of the optimality for the models for both types of responses. Efficient design construction algorithms are developed to construct the local D-and A-optimal designs for given parameter values. To capture a correlation between the two types of responses, we propose a Bayesian hierarchical modeling framework to jointly model a continuous and a binary response. Compared with the existing methods, the Bayesian method overcomes two restrictions. First, it solves the problem in which the model size (specifically, the number of parameters to be estimated) exceeds the number of observations for the continuous response. Second, the Bayesian model can provide statistical inference on the estimated parameters and predictions. Gibbs sampling scheme is used to generate accurate estimation and prediction for the Bayesian hierarchical model. Both simulation and real case study are shown to illustrate the proposed method.

Feb. 28, 2018

Enhanced Understanding of MCPMod in Dose-ranging Studies

Li Wang : 4 p.m. in SEO 636
Abstract In dose ranging clinical trials, it is critical to investigate the dose-response profile and to identify a minimum effective dose (MED) to guide the dose selection for phase 3 confirmatory trials. Traditional dose ranging trials focus on pairwise comparisons between placebo and each investigational dose, while in recent years MCP-Mod (Multiple Comparison Procedures & Modeling) arose and gained popularity in the design and analysis of dose ranging trials. Comprehensive comparison between MCP-Mod and other methods have been made on continuous variables assuming a normal distribution. We extend the comparison to binary/binomial response variables. Via simulation, the rate of correct and incorrect MED identification are compared for Dunnett's test, trend test and MCP-Mod for a variety of underlying dose response profiles including both monotone and non-monotone dose responses and are compared under a large number of trial design settings. The precision of MED estimation using MCP-Mod is also evaluated comparing the design options of more dose levels and smaller sample size per dose versus fewer dose levels and larger sample size per dose

March 7, 2018

A method for computing transition pathways of conformational changes of a biomolecule

Ruijun Zhao : 4 p.m. in SEO 636
Abstract Molecular dynamics (MD) is a computer simulation method for studying the physical movement of atoms and molecules. It has broad applications in many fields of sciences. In this talk, I will discuss how we use MD to compute transition pathways of conformational change, given two different metastable states of a biomolecule. In particular, we proposed an efficient algorithm, Maximum Flux Transition Paths, to compute such a path and applied the method to the Src tyrosine kinase family, which has long been implicated in the development of cancer. Maximum Flux Transition Paths relies on efficiently computing the free energy and proto-diffusion tensor, which are computed as the conditional expectations of some ``observable" A(x) that depend on random states x drawn from distributions that are known except for their normalizing factor. Vast amounts of computer time are used to compute these expectations. Markov chain Monte Carlo methods are very popular for computing these expectations. In this talk, I will also discuss the challenge of this method and how to estimate the accuracy of these expectations.

March 14, 2018

On I-Optimal Designs for Generalized Linear Models: An Efficient Algorithm via General Equivalence Theory

Yiou Li : 4 p.m. in SEO 636
Abstract The generalized linear model plays an important role in statistical analysis and the related design issues are undoubtedly challenging. The state-of-the-art works mostly apply to design criteria on the estimates of regression coefficients. It is of importance to study optimal designs for generalized linear models, especially on the prediction aspects. In this talk, I will discuss a prediction-oriented design criterion, I-optimality, and how we develop an efficient sequential algorithm of constructing I-optimal designs for generalized linear models. Through establishing the General Equivalence Theorem of the I-optimality for generalized linear models, an insightful understanding is obtained for the proposed algorithm on how to sequentially choose the support points and update the weights of support points of the design. The proposed algorithm is computationally efficient with guaranteed convergence property. Numerical examples are conducted to evaluate the feasibility and computational efficiency of the proposed algorithm.

March 21, 2018

Sampling for Conditional Inference on Network Data

Yuguo Chen : 4 p.m. in SEO 636
Abstract Random graphs with given vertex degrees have been widely used as a model for many real-world complex networks. We describe a sequential sampling method for sampling networks with a given degree sequence. These samples can be used to approximate closely the null distributions of a number of test statistics involved in such networks, and provide an accurate estimate of the total number of networks with given vertex degrees. We apply our method to a range of examples to demonstrate its efficiency in real problems.

April 4, 2018

Understanding the Effects of Predictor Variables in Black Box Supervised Learning Models

Daniel W. Apley : 4 p.m. in SEO 636
Abstract For many supervised learning applications, understanding and visualizing the effects of the predictor variables on the predicted response is of paramount importance. A shortcoming of black box supervised learning models (e.g., complex trees, neural networks, boosted trees, random forests, nearest neighbors, local kernel-weighted methods, support vector regression, etc.) in this regard is their lack of interpretability or transparency. Partial dependence (PD) plots, which are the most popular general approach for visualizing the effects of the predictors with black box supervised learning models, can produce erroneous results if the predictors are strongly correlated, because they require extrapolation of the response at predictor values that are far outside the multivariate envelope of the training data. Functional ANOVA for correlated inputs can avoid this extrapolation but involves prohibitive computational expense and subjective choice of additive surrogate model to fit to the supervised learning model. We present a new visualization approach that we term accumulated local effects (ALE) plots, which have a number of advantages over existing methods. First, ALE plots do not require unreliable extrapolation with correlated predictors. Second, they are orders of magnitude less computationally expensive than PD plots, and many orders of magnitude less expensive than functional ANOVA. Third, they yield convenient variable importance/sensitivity measures that possess a number of desirable properties for quantifying the impact of each predictor.

April 11, 2018

More powerful test procedures for multiple hypothesis testing

Shunpu Zhang : 4 p.m. in SEO 636
Abstract We propose a new multiple test called the minPOP test and two of its modified versions (the left truncated and the double truncated minPOP tests) for testing multiple hypotheses simultaneously. We show that these tests have multiple testing procedures based on these tests have strong control of the family-wise error rate. A method for finding the p-values of the proposed multiple testing procedures after adjusting for multiplicity is also developed. Simulation results show that the minPOP tests in general have higher global power than the existing well known multiple tests, especially when the number of hypotheses being compared is relatively large. Among the multiple testing procedures we developed, we find that the ones based on the left truncated and double truncated minPOP tests tend to have higher number of rejections than the existing multiple testing procedures. In the case of correlated test statistics, simulation results show that only the double truncated minPOP test is reasonably robust to positively correlated test statistics, while all the other tests seem to be robust to negatively correlated test statistics.

April 18, 2018

Optimal Portfolio under Fractional Stochastic Environment

Jean-Pierre Fouque : 3 p.m. in SEO 636
Abstract Rough stochastic volatility models have attracted a lot of attention recently, in particular for the linear option pricing problem. In this talk, starting with power utilities, we propose to use a martingale distortion representation of the optimal value function for the nonlinear asset allocation problem in a (non-Markovian) fractional stochastic environment (for all Hurst index $H \in (0, 1)$). We rigorously establish a first order approximation of the optimal value, when the return and volatility of the underlying asset are functions of a stationary slowly varying fractional Ornstein-Uhlenbeck process. We prove that this approximation can be also generated by the zeroth order trading strategy providing an explicit strategy which is asymptotically optimal in all admissible controls. Furthermore, we extend the discussion to general utility functions, and obtain the asymptotic optimality of this strategy in a specific family of admissible strategies. If time permits, we will also discuss the problem under fast mean-reverting fractional stochastic environment. Joint work with Ruimeng Hu (UCSB).

Counting Process Based Dimension Reduction Methods for Censored Outcomes

Ruoqing Zhu : 4 p.m. in SEO 636
Abstract We propose a class of dimension reduction methods for right censored survival data using a counting process representation of the failure process. Semiparametric estimating equations are constructed to estimate the dimension reduction subspace for the failure time model. The proposed method addresses two fundamental limitations of existing approaches. First, using the counting process formulation, it does not require any estimation of the censoring distribution to compensate the bias in estimating the dimension reduction subspace. Second, the nonparametric part in the estimating equations is adaptive to the structural dimension, hence the approach circumvents the curse of dimensionality. Asymptotic normality is established for the obtained estimators. We further propose a computationally efficient approach that requires only a singular value decomposition to estimate the dimension reduction subspace. Numerical studies suggest that the proposed methods exhibit significantly improved performance for estimating the true dimension reduction subspace. We further conducted a real data analysis on a skin cutaneous melanoma dataset from The Cancer Genome Atlas. The findings have important biological implications. The proposed methods are implemented in the R package ``orthoDr'', which efficiently solves the semiparametric estimating equations within the Stiefel manifold of the parameter space.

April 25, 2018

Concordance-Assisted Learning for Individualized Treatment Regimes

Rui Song : 4 p.m. in SEO 636
Abstract In the first part of the talk, we propose a new concordance-assisted learning for estimating optimal individualized treatment regimes. We first introduce a type of concordance function for prescribing treatment and propose a robust rank regression method for estimating the concordance function. We then find treatment regimes, up to a threshold, to maximize the concordance function, named prescriptive index. Finally, within the class of treatment regimes that maximize the concordance function, we find the optimal threshold to maximize the value function. Although this method makes better use of the available information through pairwise comparison, the objective function is discontinuous and computationally hard to optimize. In the second part of the talk, we consider a convex surrogate loss function to solve this problem. In addition, our algorithm ensures sparsity of decision rule and makes it easy to interpret. Simulation results of various settings and application to STAR*D both illustrate that the proposed method can still estimate optimal treatment regime successfully when the numb of covariates is large.

May 2, 2018

My (Mis)Adventures in Modeling and Simulation

Peter Bonate : 3 p.m. in SEO 636
Abstract Dr. Peter Bonate has over 20 years experience in modeling and simulation in the pharmaceutical industry. Dr. Bonate will discuss his career and the role modeling and simulation has played in the development of many different pharmaceutical products.

Aug. 29, 2018

Distributions of pattern statistics in sparse Markov models

Donald E.K. Martin : 4 p.m. in 636 SEO
Abstract Higher-order Markov models provide a good approximation to probabilities associated with many categorical time series, and thus they are applied extensively. However, a major drawback associated with them is that the number of model parameters grows exponentially in the order of the model, and thus only very low-order models are considered in applications. Another drawback is lack of flexibility, in that higher-order Markov models give relatively few choices for the number of model parameters. Sparse Markov models are Markov models where transition probabilities are lumped into classes comprised of invariant probabilities. The contexts for conditioning may be either hierarchical (as in variable length Markov chains) or non-hierarchical. This supplies a model that helps with the two problems given above, and which thus gives a better handling of the trade-off between bias associated with having too few model parameters and variance associated with having too many. In this work, methods for efficient computation of pattern distributions through Markov chains with minimal state spaces are extended to the sparse Markov framework.

Sept. 5, 2018

Dynamic Tensor Clustering with Applications in Neuroimaging and Online Advertising

Wei Sun : 4 p.m. in 636 SEO
Abstract Tensor as a multi-dimensional generalization of matrix has received increasing attention due to its success in many empirical tasks. In particular, dynamic tensor data are becoming prevalent since time is often one of the tensor modes. Existing tensor clustering methods either fail to account for the dynamic nature of the data, or are inapplicable to a general-order tensor. Also there is often a gap between statistical guarantee and computational efficiency for existing tensor clustering solutions. In this talk, I will introduce a new dynamic tensor clustering method, which takes into account both sparsity and fusion structures, and enjoys strong statistical guarantees as well as high computational efficiency. The efficacy of our approach will be illustrated via two real applications: brain dynamic functional connectivity analysis, and online advertisement clustering for market segmentation.

Sept. 12, 2018

Organizational meeting

Organizational meeting : 4 p.m. in 636 SEO

Sept. 19, 2018

Bridging Discrete-Time and Continuous-Time Modeling for Stochastic First-Order Optimization

Shuwen Lou : 4 p.m. in 636 SEO
Abstract We now live in a world surrounded by data. As an example, when we want to buy or sell a house, we browse real estate websites and go through related listings. By comparing the "data", we subconsciously ``generate a price quote" for the house we are interested in buying or trying to sell. This can be viewed as an optimization problem which, in theory, can be solved using gradient descent (GD) method. However, in real-world scenarios, because of the tremendous sizes of the datasets, vanilla GD is typically not an efficient or computable option. An improved version of vanilla GD is stochastic gradient descent (SGD). In the first half of this talk, we will go through the background of GD and SGD algorithms. From there, we will introduce how a discrete-time SGD algorithm can be modeled by a continuous-time stochastic process based on Brownian motion. Using probabilistic tools, one can reveal many interesting properties from continuous-time versions of SGD. Some of these properties can be translated back to discrete-time SGD algorithms.

Sept. 26, 2018

JMP and the Predictive Modeling Workflow

Kevin Potcner : 4 p.m. in 636 SEO
Abstract As the size and sources of data becomes more available in today's business environments, data analysts are beginning to add more sophisticated predictive statistical modeling techniques to their analysis toolkit. A typical real-world predictive modeling workflow includes data cleaning and exploration, model fitting, model validation, model comparison, final model selection and deployment of the final predictive model. In this presentation, a statistical scientist from JMP will illustrate the predictive modeling workflow by analyzing a real dataset. After data preparation and initial exploration, we will create a number of predictive models such as Multiple Linear Regression, Regression tree, Neural Net, and K-Nearest Neighbors. We will evaluate each model and select the best model using the Prediction Profiler and JMP's Model Comparison tool. Code will be automatically created in a variety of programming languages (e.g., SAS, SQL, Python, et al.) in order to implement that model in a production environment.

Oct. 3, 2018

Constructing Stabilized Dynamic Treatment Regimes

Guanhua Chen : 4 p.m. in 636 SEO
Abstract We propose a new method termed stabilized O-learning for deriving stabilized dynamic treatment regimes (DTRs), which are sequential decision rules for individual patients not only adapt over the course of the disease progression but also consistent over time in its format. The method provides a robust and efficient learning framework for constructing DTRs by directly optimizing a doubly robust estimator of the expected long-term outcome. It can accommodate various types of outcomes, including continuous, categorical and potentially censored survival outcomes. In addition, the method is flexible to incorporate clinical preferences into a qualitatively fixed rule, where the parameters indexing the decision rules that are shared across stages can be estimated simultaneously. We conducted extensive simulation studies, showing a superior performance of the proposed method. We analyzed the data from the prospective Canary Prostate Cancer Active Surveillance study using the proposed method.

Oct. 10, 2018

One-point function and natural parametrization for loop-erased random walk in three dimensions

Xinyi Li : 4 p.m. in 636 SEO
Abstract We consider loop-erased random walk (LERW) in three dimensions and give an asymptotic estimate on the one-point function for LERW and the non- intersection probability of LERW and simple random walk in three dimensions. Then we show that 3D LERW converges to its scaling limit in natural parametrization. This is a joint work in progress with Daisuke Shiraishi (Kyoto).

Oct. 17, 2018

Modeling Non-stationary Multivariate Time Series of Counts via Common Factors

Fangfang Wang : 4 p.m. in 636 SEO
Abstract In this talk, a new parameter-driven model for multivariate time series of counts is discussed. The time series is not necessarily stationary. The mean process is modelled as the product of modulating factors and unobserved stationary processes. The former characterizes the long-run movement in the data, while the latter is responsible for rapid fluctuations and other unknown or unavailable covariates. The unobserved stationary processes evolve independently of the past observed counts, and might interact with each other. We express the multivariate unobserved stationary processes as a linear combination of possibly low-dimensional factors that govern the contemporaneous and serial correlation within and across the observed counts. Regression coefficients in the modulating factors are estimated via pseudo maximum likelihood estimation, and identification of common factor(s) is carried out through eigenanalysis on a positive definite matrix that pertains to the autocovariance of the observed counts at nonzero lags. Theoretical validity of the two-step estimation procedure is presented. We also provide numerical results that corroborate the theoretical findings. Finally, we illustrate the use of the proposed model through an application to the numbers of National Science Foundation funding awarded to seven research universities from January 2001 to December 2012.

Oct. 24, 2018

Estimation and Inference for Differential Networks

Mladen Kolar : 4 p.m. in 636 SEO
Abstract We present a recent line of work on estimating differential networks and conducting statistical inference about parameters in a high-dimensional setting. First, we consider a Gaussian setting and show how to directly learn the difference between the graph structures. A debiasing procedure will be presented for construction of an asymptotically normal estimator of the difference. Next, building on the first part, we show how to learn the difference between two graphical models with latent variables. Linear convergence rate is established for an alternating gradient descent procedure with correct initialization. Simulation studies illustrate performance of the procedure. We also illustrate the procedure on an application in neuroscience. Finally, we will discuss how to do statistical inference on the differential networks when data are not Gaussian.

Oct. 31, 2018

Applying Mathematics and Statistics to Characterize the Effectiveness of a Pharmaceutical Product

Yi-Lin Chiu : 4 p.m. in 636 SEO
Abstract This presentation addresses the basic science for clinical pharmacology: effective and safe drug administration. We use basic mathematics and statistics to introduce the applications to pharmacology, including characterizing the drug concentration, dose selection, and dosing strategy. The utility of clinical pharmacology will be explained: We will show how drugs work, rather than asking the audience to memorize information about individual drugs. Therefore, we can understand why drugs are given, as well as when they should be given, and come up with a better way to improve the effectiveness of drug administrations.

Nov. 7, 2018

MULTILAYER TENSOR FACTORIZATION WITH APPLICATIONS TO RECOMMENDER SYSTEMS

Annie Qu : 4 p.m. in 636 SEO
Abstract Recommender systems have been widely adopted by electronic commerce and entertainment industries for individualized prediction and recommendation, which benefit consumers and improve business intelligence. In this article, we propose an innovative method, namely the recommendation engine of multilayers (REM), for tensor recommender systems. The proposed method utilizes the structure of a tensor response to integrate information from multiple modes, and creates an additional layer of nested latent factors to accommodate between-subjects dependency. One major advantage is that the proposed method is able to address the “cold-start" issue in the absence of information from new customers, new products or new contexts. Specifically, it provides more effective recommendations through sub-group information. To achieve scalable computation, we develop a new algorithm for the proposed method, which incorporates a maximum block improvement strategy into the cyclic block-wise-coordinate-descent algorithm. In theory, we investigate both algorithmic properties for global and local convergence, along with the asymptotic consistency of estimated parameters. Finally, the proposed method is applied in simulations and IRI marketing data with 116 million observations of product sales. Numerical studies demonstrate that the proposed method outperforms existing competitors in the literature. This is joint work with Xuan Bi and Xiaotong Shen.

Nov. 14, 2018

Factorizations and estimates of Dirichlet heat kernels for non-local operators with critical killings

Renming Song : 4 p.m. in 636 SEO
Abstract In this talk I will discuss heat kernel estimates for critical perturbations of non-local operators. To be more precise, let $X$ be the reflected $\alpha$-stable process in the closure of a smooth open set $D$, and $X^D$ the process killed upon exiting $D$. We consider potentials of the form $\kappa(x)=C\delta_D(x)^{-\alpha}$ with positive $C$ and the corresponding Feynman-Kac semigroups. Such potentials do not belong to the Kato class. We obtain sharp two-sided estimates for the heat kernel of the perturbed semigroups. The interior estimates of the heat kernels have the usual $\alpha$-stable form, while the boundary decay is of the form $\delta_D(x)^p$ with non-negative $p\in [\alpha-1, \alpha)$ depending on the precise value of the constant $C$. Our result recovers the heat kernel estimates of both the censored and the killed stable process in $D$. Analogous estimates are obtained for the heat kernel of the Feynman-Kac semigroup of the $\alpha$-stable process in ${\mathbf R}^d\setminus \{0\}$ through the potential $C|x|^{-\alpha}$. All estimates are derived from a more general result described as follows: Let $X$ be a Hunt process on a locally compact separable metric space in a strong duality with $\widehat{X}$. Assume that transition densities of $X$ and $\widehat{X}$ are comparable to the function $\widetilde{q}(t,x,y)$ defined in terms of the volume of balls and a certain scaling function. For an open set $D$ consider the killed process $X^D$, and a critical smooth measure on $D$ with the corresponding positive additive functional $(A_t)$. We show that the heat kernel of the the Feynman-Kac semigroup of $X^D$ through the multiplicative functional $\exp(-A_t)$ admits the factorization of the form ${\mathbf P}_x(\zeta >t)\widehat{\mathbf P}_y(\widehat{\zeta}>t)\widetilde{q}(t,x,y)$. This is joint work with Soobin Cho, Panki Kim and Zoran Vondracek.

Nov. 28, 2018

Introduction of Actuarial Profession

Linyi Zhang : 4 p.m. in 636 SEO

Feb. 20, 2019

Posterior Contraction and Credible Sets for Filaments of Regression Functions

Subhashis Ghoshal : 4 p.m. in 636 SEO
Abstract The filament of a smooth function f consists of local maximizers of f when moving in a certain direction. The filament is an important geometrical feature of the surface of the graph of a function. It is also considered as an important lower dimensional summary in analyzing multivariate data. There have been some recent theoretical studies on estimating filaments of a density function using a nonparametric kernel density estimator. In this talk, we consider a Bayesian approach and concentrate on the nonparametric regression problem. We study the posterior contraction rates for filaments using a finite random series of B-splines prior on the regression function. Compared with the kernel method, this has the advantage that the bias can be better controlled when the function is smoother, which allows obtaining better rates. Under an isotropic Holder smoothness condition, we obtain the posterior contraction rate for the filament under two different metrics --- a distance of separation along an integral curve, and the Hausdorff distance between sets. Moreover, we construct credible sets of optimal size for the filament with sufficient frequentist coverage. We study the performance of our proposed method through a simulation study and apply on a dataset on California earthquakes to assess the fault-line of the maximum local earthquake intensity. Based on joint work with my former graduate student, Dr. Wei Li, Assistant Professor, Syracuse University, New York.

Feb. 27, 2019

Using Prior Information for Intelligent Factor Allocation and Design Selection

William Li : 4 p.m. in 636 SEO
Abstract While literature on constructing efficient experimental designs has been plentiful, how best to incorporate prior information when assigning factors to the columns has received little attention. This talk summarizes a series of recent studies that focus on information of individual columns. For regular designs, we propose the individual word length pattern (iWLP) that can be used to rank columns. With prior information on how likely a factor is important, iWLP can be used to intelligently assign factors to columns, and select the best designs to accommodate such prior information. This criterion is then extended to study nonregular designs, which we denote as the individual generalized word length pattern (iGWLP). We illustrate how iGWLP helps to identify important differences in the aliasing that is likely otherwise missed. Given the complexity of characterizing partial aliasing, iGWLP will help practitioners make more informed assignment of factors to columns when utilizing nonregular fractions. The theoretical justifications of the proposed iGWLP are provided in terms of statistical model and projection properties. In the third part, we consider clear effects involving an individual column (iCE). Motivated by a real application, we introduce the clear effects pattern, derived from iCE, and propose a class of designs called maximized clear effects pattern (MCEP) designs. We compare MCEP designs with commonly used minimum aberration designs and MaxC2 designs that maximize the number of clear two-factor interaction. We also extend the definition of iCE and MCEP designs by considering blocking schemes.

March 6, 2019

A Super Scalable Algorithm for Short Segment Detection

Yue Niu : 4 p.m. in 636 SEO
Abstract In many applications such as copy number variant (CNV) detection, the goal is to identify short segments on which the observations have different means or medians from the background. Those segments are usually short and hidden in a long sequence, and hence are very challenging to find. We study a super scalable short segment (4S) detection algorithm in this paper. This nonparametric method clusters the locations where the observations exceed a threshold for segment detection. It is computationally efficient and does not rely on Gaussian noise assumption. Moreover, we develop a framework to assign significance levels for detected segments. We demonstrate the advantages of our proposed method by theoretical, simulation, and real data studies.

March 13, 2019

Identifying Appropriate Probabilistic Models for Sparse Discrete Data

Hani Aldirawi : 4 p.m. in 636 SEO
Abstract Modeling sparse and discrete data such as microbiome and insurance claim data is challenging due to the exceeded number of zeros. Many probabilistic models have been used for modeling sparse data, including Poisson, negative binomial, zero-inflated Poisson, and zero-inflated negative binomial models. We propose a statistical procedure for identifying the most appropriate discrete probabilistic models for zero-inflated or Hurdle models based on the p-value of the discrete Kolmogorov-Smirnov (KS) test when the population parameters are unknown. We develop a general procedure for estimating the parameters for a large class of zero-inflated models and Hurdle models. We also develop a general likelihood ratio test based on Neyman-Pearson lemma for choosing the best model when appropriate ones are more than one.

March 20, 2019

Finite sample change point inference and identification for high-dimensional location-shift

Mengjia Yu : 4 p.m. in 636 SEO
Abstract I will discuss two approaches of the cumulative sum (CUSUM) statistics and the U-statistics in change point problems for high-dimensional location-shift. Both works are non-parametric, fully data-dependent and enjoying strong theoretical guarantees under arbitrary dependence structures. 1. Based on the $\ell^{\infty}$-norm of the CUSUM statistics, we study inference and identification for high-dimensional mean vectors. For the problem of testing existence of a change point in an independent sample generated from the mean-shift model, we introduce a Gaussian multiplier bootstrap to calibrate critical values of the CUSUM test statistics. For the problem of estimating the change point location once it is detected, two estimators are proposed by maximizing the $\ell^{\infty}$-norm of the generalized CUSUM statistics at two different weighting scales. In both problems, dimension impacts the rate of convergence only through the logarithm factors, and therefore consistency of the CUSUM location estimators is possible when $p$ is much larger than $n$. 2. In cases where mean does not exist, we consider signal cancellations in the general U-statistics framework with anti-symmetric kernels of order 2, and proposed another test that is more robust to detect location-shift by selecting bounded kernels. The $\ell^{\infty}$-norm of the U-statistic and its Gaussian multiplier bootstrap approximation are investigated, and no tuning parameter is needed in this scheme. Subject to mild conditions kernels, we derive similar rates of uniform convergence to our CUSUM-based test. Connection of two approaches and numeric studies are also provided.

April 3, 2019

Dynamic Multiscale Spatiotemporal Models for Multivariate Gaussian Data

Marco Ferreira : 4 p.m. in 636 SEO
Abstract We discuss classes of dynamic multiscale models for multivariate Gaussian spatiotemporal data. First, we develop multiscale spatial factorizations to decompose the data at each time point into spatiotemporal multiscale coefficients. We then connect these spatiotemporal multiscale coefficients through time with state-space evolutions. Further, we propose simulation-based Bayesian posterior analysis. In particular, we develop filtering equations for updating of information forward in time and smoothing equations for integration of information backward in time, and use these equations to develop forward filter backward samplers for the spatiotemporal multiscale coefficients. Because the multiscale coefficients are conditionally independent a posteriori, our Bayesian posterior analysis is scalable, computationally efficient, and highly parallelizable. Finally, we illustrate the usefulness of our dynamic multiscale spatiotemporal methodology with applications to multivariate spatiotemporal data on temperatures in the upper troposphere and lower stratosphere over North America.

April 10, 2019

Overview of ICH E9(R1): Estimands and Sensitivity Analysis in Clinical Trials

Jane Qian : 4 p.m. in 636 SEO
Abstract ICH is the International Council for Harmonization of Technical Requirements for Pharmaceuticals for Human Use. ICH E9 guideline (“Statistical Principles for Clinical Trials”) was published in 1995 and has since served as a foundation of regulatory guidance on major statistical aspects of confirmatory clinical trials. In October 2014, the Steering Committee of ICH endorsed the formation of an expert working group to develop an addendum to the ICH E9 guideline. In 2017, the addendum ICH E9(R1) (Estimands and sensitivity analyses in clinical trials) was released for public comment. ICH E9(R1) focus on two topics involving randomized confirmatory clinical trials: estimands and sensitivity analyses. Both topics are motivated by the need to improve the precision with which scientific questions of interest are formulated and addressed by clinical researchers and regulators, specifically in the context of post-randomization (intercurrent) events such as use of rescue medication or missing data. In this seminar, an overview of ICH E9(R1) will be given with the focus on why it was necessary to develop this ICH E9 addendum, what it entails and the impact of the addendum on future clinical trial design and analyses.

April 17, 2019

Spatially Dependent Functional Data: Covariance Estimation, Principal Component Analysis, and Kriging

Yehua Li : 4 p.m. in 636 SEO
Abstract We consider spatially dependent functional data collected under a geostatistics setting, where locations are sampled from a spatial point process and a random function is observed at each location. The functional response is the sum of a spatially dependent functional effect and a spatially independent functional nugget effect. Observations on each function are made on discrete time points and contaminated with measurement errors. Under the assumption of spatial stationarity and isotropy, we propose a tensor product spline estimator for the spatio-temporal covariance function. If a coregionalization covariance structure is further assumed, we propose a new functional principal component analysis method that borrows information from neighboring functions. Under a unified framework for both sparse and dense functional data, where the number of observations per curve is allowed to be of any rate relative to the number of functions, we develop the asymptotic convergence rates for the proposed estimators. Advantages of the proposed approach over existing methods are demonstrated through simulation studies and a real data application to the home price-rent ratio data in the San Francisco Bay Area.

A smooth collaborative recommender system

Junhui Wang : 3 p.m. in 636 SEO
Abstract In recent years, there has been a growing demand to develop efficient recommender systems which track users' preferences and recommend potential items of interest to users. In this talk, I will present a smooth collaborative recommender system to utilize dependency information among users and items which share similar characteristics under the singular value decomposition framework. The proposed method incorporates the neighborhood structure among user-item pairs by exploiting covariates to improve the prediction performance. One key advantage of the proposed method is that it leads to more efficient recommendation for "cold-start" users and items, whose preference information is completely missing from the training set. As this type of data involves large-scale customer records, efficient scheme will be proposed to achieve scalable computing. The advantage is confirmed in a variety of simulated experiments as well as one large-scale real example on <i>Last.fm</i> music listening counts. If time permits, the asymptotic properties will also be discussed.

April 24, 2019

Sufficient dimension folding via distance covariance

Wenhui Sheng : 4 p.m. in 636 SEO
Abstract We propose a new sufficient dimension folding method using distance covariance for regression in which the predictors are matrix- or array-valued. The method works efficiently without strict assumptions on the predictor. It is modelfree and neither smoothing techniques or selection of tuning parameters is needed. Moreover, it works for both univariate and multivariate response cases. We use two approaches to estimate the structural dimensions: bootstrap method and a new method of local search. Simulations and real data analysis support the efficiency and effectiveness of the method.

May 1, 2019

Weak Dependence Conditions for High-Dimensional Inference: Applications to Group Comparisons

Solomon Harrar : 4 p.m. in 636 SEO
Abstract Recent results for high-dimensional inference make assumptions that require weak dependence (pseudo independence) between the variables. These requirements fail to be satisfied, for example, for all elliptically contoured distributions except for normal distribution. In this talk, we present weaker dependence conditions for high-dimensional asymptotic theory. With these conditions the scope of application of many high-dimensional results broadens substantially. For example, mixing-type dependence and general conditions on variance of quadratic forms are covered. The application of the new conditions will be demonstrated with high-dimensional tests for comparing group differences in terms of means and in terms of Mann-Whitney effects. The later is particularly useful for non-metric data such as ordered categorical data, and also for skewed and heavy tailed continuous data. Simulation results show favorable performance of these tests. Data from Electroencephalograph (EEG) experiment is analyzed to illustrate these applications. The results presented in this talk are joint works with Xiaoli Kong, Department of Mathematics and Statistics, Loyola University-Chicago

May 13, 2019

EzGP: Easy-to-Interpret Gaussian Process Models for Computer Experiments with Both Quantitative and Qualitative Factors

Abhyuday Mandal : 3 p.m. in 636 SEO
Abstract Computer experiments with both quantitative and qualitative inputs are commonly used in science and engineering applications. Constructing desirable emulators for such computer experiments remains a challenging problem. Here we propose an easy-to-interpret Gaussian process (EzGP) model for computer experiments to reflect the change of the computer model under different level combinations of qualitative factors. The proposed modeling strategy, based on an additive Gaussian process, is flexible to address the heterogeneity of computer models involving multiple qualitative factors. We also develop two useful variants of the EzGP model to achieve computation efficiency when dealing with high dimensional data and large data size. The merits of these models are illustrated by a real data application and several numerical examples.

A Class of New Particle Swarm Optimization Algorithms and Applications to Model Discriminating Designs

William Li : 4 p.m. in 636 SEO
Abstract Exchange-type of algorithms have been commonly used in design construction problems. In recent years, algorithms based on Particle Swarm Optimization (PSO) techniques have been proposed to construct optimal designs. PSO algorithms have been developed mostly for continuous-type of problems. In this talk we develop a general class of Particle Swarm Exchange algorithms, targeting discrete-type of design problems . We used the proposed algorithm to construct a class of optimal model-discriminating designs. It is shown that the algorithms work both efficiently and effectively. The algorithm compared favorably with the coordinate-exchange algorithm - one of the most commonly used algorithms. And we obtained model-discriminating designs that are comparable and sometime better than existing results.

Aug. 28, 2019

Organizational meeting

Yichao Wu : 4 p.m. in 636 SEO

Sept. 4, 2019

Envelope-based Sparse Partial Least Squares

Zhihua Su : 4 p.m. in 636 SEO
Abstract Sparse partial least squares (SPLS) is widely used in applied sciences as a method that performs dimension reduction and variable selection simultaneously in linear regression. Several implementations of SPLS have been derived, among which the SPLS proposed in Chun and Keleş (2010) is very popular and highly cited. However, for all of these implementations, the theoretical properties of SPLS are largely unknown. In this paper, we propose a new version of SPLS, called the envelope-based SPLS, using a connection between envelope models and partial least squares (PLS). We establish the consistency, oracle property and asymptotic normality of the envelope-based SPLS estimator. The large-sample scenario and high-dimensional scenario are both considered. We also develop the envelope-based SPLS estimators under the context of generalized linear models, and discuss its theoretical properties including consistency, oracle property and asymptotic distribution. Numerical experiments and examples show that the envelope-based SPLS estimator has better variable selection and prediction performance over the existing SPLS estimators.

Sept. 11, 2019

Pharmacometrics: Application of MSCS to Pharmaceutical Development

Stacey Tannenbaum : 4 p.m. in 636 SEO
Abstract Pharmacometrics is the application of biological and pharmacological science and statistical/ mathematical/ computational methods to optimize pharmaceutical development. The goal of a pharmacometrician is to get the right dose of the right drug to the right patient at the right time! Pharmacometrics includes a wide span of models and applications, but the primary focus of the seminar is on understanding the pharmacokinetics of the drug (how a drug is absorbed, processed, distributed, and eliminated, and how much is in the plasma and site of action at a given time). Every subject will have their own unique concentration-time profile for a given dose and formulation, which is dependent upon intrinsic and extrinsic factors (weight, smoking status, other drugs, health state, etc). One important job of the pharmacometrician is to understand the quantitative impact of these factors on the pharmacokinetics, and to determine which have enough of an impact to make a change to the recommended dose. Once a model is fit to a patient population, simulations can be performed to assess the outcome with different inputs (higher and lower doses, less/more frequent doses, etc). The seminar will also include a discussion of some of the technical details of Pharmacometrics, including some of the mathematical methods, software, data sources, validation techniques, and challenges that pharmacometricians face in their day-to-day work.

Sept. 18, 2019

A sparse clustering algorithm for identifying cluster changes across conditions with applications in single-cell RNA-sequencing data

Jun Li : 4 p.m. in 636 SEO
Abstract Clustering analysis, in its traditional setting, identifies groupings of samples from a single population/condition. We consider a different setting when the data available are samples from two different conditions, such as cells before and after drug treatment. Cell types in cell populations change as the condition changes: some cell types die out, new cell types may emerge, and surviving cell types evolve to adapt to the new condition. Using single-cell RNA-sequencing data that measure the gene expression of cells before and after the condition change, we propose an algorithm, SparseDC, which identifies cell types, traces their changes across conditions, and identifies genes which are marker genes for these changes. By solving a unified optimization problem, SparseDC completes all three tasks simultaneously. As a general algorithm that detects shared/distinct clusters for two groups of samples, SparseDC can be applied to problems outside the field of biology.

Sept. 25, 2019

Random Forest Prediction Intervals

Dan Nettleton : 4 p.m. in 636 SEO
Abstract Breiman's seminal paper on random forests has more than 30,000 citations according to Google Scholar. The impact of Breiman's random forests on machine learning, data analysis, data science, and science in general is difficult to measure but unquestionably substantial. The virtues of random forest methodology include no need to specify functional forms relating predictors to a response variable, capable performance for low-sample-size high-dimensional data, general prediction accuracy, easy parallelization, few tuning parameters, and applicability to a wide range of prediction problems with categorical or continuous responses. Like many algorithmic approaches to prediction, random forests are typically used to produce point predictions that are not accompanied by information about how far those predictions may be from true response values. From the statistical point of view, this is unacceptable; a key characteristic that distinguishes statistically rigorous approaches to prediction from others is the ability to provide quantifiably accurate assessments of prediction error from the same data used to generate point predictions. Thus, we develop a prediction interval -- based on a random forest prediction -- that gives a range of values that will contain an unknown continuous univariate response with any specified level of confidence. We illustrate our proposed approach to interval construction with examples and demonstrate its effectiveness relative to other approaches for interval construction using random forests.

Oct. 2, 2019

Equivariant Variance estimation for multiple change-point model

Ning Hao : 4 p.m. in 636 SEO
Abstract The variance of noise plays an important role in many change-point detection tools and their inference. For example, in binary segmentation or other stepwise detection methods, the variance is necessary to decide when to stop the procedure. In practice, people usually use some ad-hoc methods to estimate the noise variance. However, these methods may be problematic when there are many change points. We will introduce an equivariant variance estimator and show its advantages over existing methods. This talk is based on a joint work with Yue S. Niu and Han Xiao.

Oct. 9, 2019

A Martingale Approach for Fractional Brownian Motions and Related Path Dependent PDEs

Jianfeng Zhang : 3 p.m. in 636 SEO
Abstract Motivated by option pricing in a financial market with rough volatility, we study backward SDEs in a framework where the (forward) state process satisfies a Volterra type SDE, with fractional Brownian motion as a typical example. Such processes are neither Markov processes nor semimartingales, and most notably, they feature a certain time inconsistency which makes any direct application of Markovian ideas impossible without passing to a path-dependent framework. Our main result is a functional Ito formula, extending the seminal work of Dupire to our more general framework. In particular, unlike in Dupire's setting where one needs only to consider the stopped paths, here we need to concatenate the observed path up to the current time with a certain smooth observable curve derived from the distribution of the future paths. This new feature is due to the time inconsistency involved in this paper. We then derive the path dependent PDEs for the backward problems. The talk is based on a joint work with Frederi Viens.

Weighted empirical minimum distance estimators in Berkson measurement error regression models

Hira Koul : 4 p.m. in 636 SEO
Abstract We develop analogs of the two classes of weighted empirical min- imum distance estimators of the underlying parameters in linear and nonlinear regression models when covariates are observed with Berk- son measurement error. One class is based on the integral of the square of symmetrized weighted empirical of residuals while the other is based on a similar integral involving a weighted empirical of residual ranks. The former class requires the regression and measurement errors to be symmetric around zero while the latter class does not need any such assumption. The first class of estimators includes the analogs of the least absolute deviation and Hodges-Lehmann estimators while the second class includes an estimator that is asymptotically more effi- cient than these two estimators at some error distributions when there is no measurement error. In the case of linear model, no knowledge of the measurement error distribution is needed. Such information is typically needed for non-linear models. We first develop these esti- mators for nonlinear models when the measurement error distribution is known and then their analogs, when this distribution is not known but validation data is available.

Oct. 16, 2019

A statistical roadmap for journey from real-world data to real-world evidence

Yixin Fang : 4 p.m. in 636 SEO
Abstract Abstract: Randomized controlled clinical trials (RCTs) are the gold standard for evaluating the safety and efficacy of pharmaceutical drugs, but in many cases their costs, duration, limited generalizability, and ethical or technical feasibility have caused some to look for real-world studies as alternatives. On the other hand, real-world data may be much less convincing due to the lack of randomization and the presence of confounding bias. In this article, we propose a statistical roadmap to translate real-world data (RWD) to robust real-world evidence (RWE). The Food and Drug Administration (FDA) is working on guidelines, with a target to release a draft by 2021, to harmonize RWD applications and monitor the safety and effectiveness of pharmaceutical drugs using RWE. The proposed roadmap aligns with the newly released framework for FDA's RWE Program in December 2018 and we hope this statistical roadmap is useful for statisticians who are eager to embark on their journeys in the real-world research.

Oct. 23, 2019

A New Framework for Distance and Kernel-based Metrics in High Dimensions

Xianyang Zhang : 4 p.m. in 636 SEO
Abstract We present new metrics to quantify and test for (i) the equality of distributions and (ii) the independence between two high-dimensional random vectors. We show that the energy distance based on the usual Euclidean distance cannot completely characterize the homogeneity of two high-dimensional distributions in the sense that it only detects the equality of means and the traces of covariance matrices in the high-dimensional setup. We propose a new class of metrics which inherit the desirable properties of the energy distance/distance covariance in the low-dimensional setting and is capable of detecting the homogeneity of/ completely characterizing independence between the low-dimensional marginal distributions in the high dimensional setup. We further propose t-tests based on the new metrics to perform high-dimensional two-sample testing/ independence testing and study its asymptotic behavior under both high dimension low sample size (HDLSS) and high dimension medium sample size (HDMSS) setups. The computational complexity of the t-tests only grows linearly with the dimension and thus is scalable to very high dimensional data. We demonstrate the superior power behavior of the proposed tests for homogeneity of distributions and independence via both simulated and real datasets.

Oct. 30, 2019

Repro Sampling Method for Joint Inference of Model Selection and Regression Coefficients in High Dimensional Linear Models

Minge Xie : 4 p.m. in 636 SEO
Abstract This paper proposes a new and effective simulation-based approach, called Repro Sampling method, to conduct statistical inference in high dimensional linear models. The Repro method creates and studies the performance of artificial samples (referred to as Repro samples) that are generated by mimicking the sampling mechanism that generated the true observed sample. By doing so, this method provides a new way to quantify model and parameter uncertainty and provide confidence sets with guaranteed coverage rates on a wide range of problems. A general theoretical framework and an effective Monte-Carlo algorithm, with supporting theories, are developed for high dimensional linear models. This method is used to jointly create confidence sets of selected models and model coefficients, with both exact and asymptotic inferences theories provided. It also provides a theoretical development to support the computational efficiency. Furthermore, this development allows us to handle inference problems involving covariates that are perfectly correlated. A new and intuitive graphical tool to present uncertainties in model selection and regression parameter estimation is also developed. We provide numerical studies to demonstrate the utility of the proposed method in a range of problems. Numerical comparisons suggest that the method is far better (in terms of improved coverage rates and significantly reduced sizes of confidence sets) than the approaches that are currently used in the literature. The development provides a simple and effective solution for the difficult post-selection inference problems.

Nov. 6, 2019

Understanding Dynamical Patterns in Complex Substitutive Systems

Dr. Ching Jin : 4 p.m. in 636 SEO
Abstract Diffusion processes are central to human interactions. One common prediction of the current modeling frameworks is that initial spreading dynamics follow exponential growth. Here we find that, for subjects ranging from mobile handsets to automobiles and from smartphone apps to scientific fields, early growth patterns follow a power law with non-integer exponents. We test the hypothesis that mechanisms specific to substitution dynamics may play a role, by analyzing unique data tracing 3.6 million individuals substituting different mobile handsets. We uncover three generic ingredients governing substitutions, allowing us to develop a minimal substitution model, which not only explains the power-law growth, but also collapses diverse growth trajectories of individual constituents into a single curve. These results offer a mechanistic understanding of power-law early growth patterns emerging from various domains and demonstrate that substitution dynamics are governed by robust self-organizing principles that go beyond the particulars of individual systems. This talk is based on my recent Nature Human Behaviour paper (attached with the email). If we have enough time, I would also like to share a couple of follow-ups of the paper or a couple of related projects we are working on recently.

Nov. 13, 2019

Unbiased Estimation and Median-Unbiasedness in Finite Population Survey Sampling

Jennifer Pajda-Delao : 4 p.m. in 636 SEO
Abstract This talk will introduce survey sampling along with some sampling designs. Then we discuss the minimum, maximum, and median as important parameters in finite population sampling. We can prove that there are no unbiased estimators of the minimum, maximum, or median for finite population sampling under any sampling design except census. We then identify and characterize a family of sampling designs such that, under these designs, the sample median is a median-unbiased estimator of the population median. In particular, we consider the simple random sampling case.

Nov. 20, 2019

Moment Kernel for Estimating Central Mean Subspace and Central Subspace

Xiangrong Yin : 4 p.m. in 636 SEO
Abstract The T-central subspace, introduced by Luo, Li and Yin (2014), allows one to perform sufficient dimension reduction for any statistical functional of interest. We propose a general estimator using (third) moment kernel to estimate the T-central subspace. In this talk, we particularly focus on central mean subspace via the regression mean function, and central subspace via Fourier transform or slicing. Theoretical results are established and simulation studies show the advantages of our proposed methods.

Dec. 4, 2019

Bayesian high-dimensional logit models: categorical responses and group sparsity

Seonghyun Jeong : 4:15 p.m. in 636 SEO
Abstract This study investigates frequentist properties of Bayesian high-dimensional logit models for categorical response variables. For high-dimensional regression coefficients, group sparse modeling is adopted to handle model selection with categorical responses. A product of a point mass and a Laplace-type distribution is used for the prior distribution on sparse regression coefficients. The procedure exhibits nearly optimal posterior contraction. A shape approximation to the posterior distribution is characterized to show model selection consistency. The distributional approximation also leads to a Bernstein-von Mises theorem for uncertainty quantification through credible sets with guaranteed frequentist coverage.

Feb. 19, 2020

Improved Shrinkage Prediction under a Spiked Covariance Structure

Trambak Banerjee : 4:15 p.m. in 636 SEO
Abstract We develop a novel shrinkage rule for prediction in a high-dimensional non-exchangeable hierarchical Gaussian model with an unknown spiked covariance structure. We propose a family of commutative priors for the mean parameter, governed by a power hyper-parameter, which encompasses from perfect independence to highly dependent scenarios. Corresponding to popular loss functions such as quadratic, generalized absolute, and linex losses, these prior models induce a wide class of shrinkage predictors that involve quadratic forms of smooth functions of the unknown covariance. By using uniformly consistent estimators of these quadratic forms, we propose an efficient procedure for evaluating these predictors which outperforms factor model based direct plug-in approaches. We further improve our predictors by introspecting possible reduction in their variability through a novel coordinate-wise shrinkage policy that only uses covariance level information and can be adaptively tuned using the sample eigen structure. We extend our methodology to aggregation based prescriptive analysis of generic multidimensional linear functionals of the predictors that arise in many contemporary applications involving forecasting decisions on portfolios or combined predictions from dis-aggregative level data. We propose an easy-to-implement functional substitution method for predicting linearly aggregative targets and establish asymptotic optimality of our proposed procedure. We present simulation experiments as well as real data examples illustrating the efficacy of the proposed method.

Feb. 26, 2020

Recent advances in statistical inference for SPDEs

Hyun-Jung Kim : 4 p.m. in 636 SEO
Abstract In this talk, we discuss recent discoveries in statistical inference for stochastic partial differential equations (SPDEs). We mainly focus on parameter estimation problems in stochastic evolution equations driven by additive noise: 1. space-time and 2. space-only colored (or white) noise. The goal of this talk is to derive "good" estimators in the sense that they are consistent and asymptotically normal to a true parameter in a specific asymptotic regime when continuous or discrete sampling of the solution process is available.

March 4, 2020

CANCELLED

Xiaotong Shen : 4 p.m. in 636 SEO

Cubature method and machine learning to solve Path Dependent PDE(PPDE)

Qi Feng : 3 p.m. in 636 SEO
Abstract The classical models for asset processes in math finance are SDEs driven by Brownian motion of the following type $X_t=x+\int_0^tb(s,X_s)ds+\int_0^t\sigma(s,X_s)\circ dB_s$. Then $u(t,X_t)=\mathbb E[{g(X_T)}|\mathcal F_{t}^X]$ is a deterministic function of $X_t$ and $u(t,x)$ solves a parabolic PDE. The cubature formula is first constructed to numerically compute functionals like $\mathbb E^{\mathbb P}[g(X_T)]$, which can be seen as a discrete approximation of the infinite dimensional Wiener measure (denoted as $\mathbb P$). In this talk, we will consider that the asset process follows a rough volatility model. For example, in the rough Heston model, the process $X_t$ is the solution of Volterra type SDEs. In this case, $X$ itself is non-Markovian, then $u(t,X_t)$ will depend on the whole path of $(X_s)_{0\le s\le t}$ and $u(t,X_{[0,t]})$ solves the so-called Path Dependent PDE (PPDE). We propose a new algorithm to numerically solve PPDE by using cubature type formulas for Volterra SDEs. The cubature formula for Volterra SDEs is solved by using machine learning method. In the end, I will show some numerical examples. The talk is based on a joint work with Jianfeng Zhang.

March 18, 2020

Cancelled

Willaim Li : 4 p.m. in 636 SEO

April 1, 2020

TBA

Ivan Nourdin : 4 p.m. in 636 SEO
Abstract TBA

April 8, 2020

CANCELLED

Lanju Zhang : 4 p.m. in 636 SEO

April 15, 2020

CANCELLED

Rina Foygel Barber : 3 p.m. in 636 SEO

April 22, 2020

CANCELLED

Yongzhao Shao : 4 p.m. in 636 SEO

April 29, 2020

CANCELLED

Peng Zeng : 4 p.m. in 636 SEO

Aug. 26, 2020

Organizational meeting

No speaker : 4 p.m. in Zoom

Sept. 2, 2020

Model-Free Variable Selection With Matrix-Valued Predictors

Yuexiao Dong : 4 p.m. in Zoom
Abstract We introduce a novel framework for model-free variable selection with matrix-valued predictors. To test the importance of rows, columns, and submatrices of the predictor matrix in terms of predicting the response, three types of hypotheses are formulated under a unified framework. The asymptotic properties of the test statistics under the null hypothesis are established and a permutation testing algorithm is also introduced to approximate the distribution of the test statistics. A maximum ratio criterion (MRC) is proposed to facilitate the model-free variable selection. Unlike the traditional stepwise regression procedures that require calculating p-values at each step, the MRC is a non-iterative procedure that does not require p-value calculation and is guaranteed to achieve variable selection consistency under mild conditions. Performance of the proposed method is evaluated in extensive simulations and demonstrated through the analysis of an electroencephalography data.

Sept. 16, 2020

Randomization Inference beyond the Sharp Null: Bounded Null Hypotheses and Quantiles of Individual Treatment Effects

Xinran Li : 4 p.m. in 636 SEO
Abstract Randomization (a.k.a. permutation) inference is typically interpreted as testing Fisher's ``sharp'' null hypothesis that all effects are exactly zero. This hypothesis is often criticized as uninteresting and implausible. We show, however, that many randomization tests are also valid for a ``bounded'' null hypothesis under which effects are all negative (or positive) for all units but otherwise heterogeneous. The bounded null is closely related to important concepts such as monotonicity and Pareto efficiency. Inverting tests of this hypothesis yields confidence intervals for the maximum (or minimum) individual treatment effect. We then extend randomization tests to infer other quantiles of individual effects, which equivalently infers proportions of units with effects larger (or smaller) than any thresholds. The proposed confidence intervals for all quantiles of individual effects are simultaneously valid, in the sense that no correction due to multiple analyses is needed. In sum, we provide a broader justification for Fisher randomization tests, and develop exact nonparametric inference for quantiles of heterogeneous individual effects. The proposed methods move beyond usual constant effects under Fisher randomization tests and average effect in Neyman's repeated sampling inference. We illustrate our methods with simulations and applications, where we find that Stephenson rank statistics often provide the most informative results.

Sept. 23, 2020

Testing goodness-of-fit and conditional independence with approximate co-sufficient sampling

Rina Foygel Barber : 4 p.m. in Zoom
Abstract Goodness-of-fit (GoF) testing is ubiquitous in statistics, with direct ties to model selection, confidence interval construction, conditional independence testing, and multiple testing, just to name a few applications. While testing the GoF of a simple (point) null hypothesis provides an analyst great flexibility in the choice of test statistic while still ensuring validity, most GoF tests for composite null hypotheses are far more constrained, as the test statistic must have a tractable distribution over the entire null model space. A notable exception is co-sufficient sampling (CSS): resampling the data conditional on a sufficient statistic for the null model guarantees valid GoF testing using any test statistic the analyst chooses. But CSS testing requires the null model to have a compact (in an information-theoretic sense) sufficient statistic, which only holds for a very limited class of models; even for a null model as simple as logistic regression, CSS testing is powerless. In this paper, we leverage the concept of approximate sufficiency to generalize CSS testing to essentially any parametric model with an asymptotically-efficient estimator; we call our extension “approximate CSS” (aCSS) testing. We quantify the finite-sample Type I error inflation of aCSS testing and show that it is vanishing under standard maximum likelihood asymptotics, for any choice of test statistic. We apply our proposed procedure both theoretically and in simulation to a number of models of interest to demonstrate its finite-sample Type I error and power. This work is joint with Lucas Janson.

Sept. 30, 2020

Statistical Inference for Geometric and Topological Data

Jisu Kim : 4 p.m. in Zoom
Abstract Geometric and topological structures can aid statistics in several ways. In high dimensional statistics, geometric structures can be used to reduce dimensionality. High dimensional data entails the curse of dimensionality, which can be avoided if there are low dimensional geometric structures. On the other hand, geometric and topological structures also provide useful information. Structures may carry scientific meaning about the data and can be used as features to enhance supervised or unsupervised learning. In this talk, I will explore how statistical inference can be done on geometric and topological structures. First, given a manifold assumption, I will explore the minimax rates of dimension estimator and reach estimator. First, given a manifold assumption, I will explore the minimax rate for estimating the dimension of the manifold. Second, also under the manifold assumption, I will explore the minimax rate for estimating the reach, which is a regularity quantity depicting how a manifold is smooth and far from self-intersecting. Third, I will investigate inference on cluster trees, which is a hierarchy tree of high-density clusters of a density function. Fourth, I will investigate inference on persistent homology of a density function, which is a representation of topological features of the density function at different levels. Third, I will present R package TDA for computing topological data analysis, which is a set of data analysis tools utilizing topology and includes persistent homology.

Oct. 7, 2020

Solving a class of time-inconsistent problems

Qingshuo Song : 4 p.m. in Zoom
Abstract The characterization of the efficient frontier in Markowitz portfolio optimization is to minimize a linear combination of mean and variance of the terminal stock price. Such a problem is known as the time-inconsistent optimization and the main difficulty is due to the failure of the dynamic programming principle. The existing approaches are game-theoretic framework and decoupling techniques on its FBSDE formulation. In this talk, we will discuss an alternative approach. The key observation is to identify the linear-quadratic structure of the underlying optimization as a function of probability distribution. This leads to explicit solutions of a class of master equations, which provides the optimal strategy to a class of time-inconsistent optimizations. Some extensions to partially observed systems will be considered briefly if time is permitted. The discussion is based on a manuscript available at https://arxiv.org/pdf/1910.05236.pdf.

Oct. 14, 2020

Graph-based change-point detection for non-Euclidean and multivariate data

Lynna Chu : 4 p.m. in Zoom
Abstract We present a new framework for the testing and estimation of change-points, locations where the distribution abruptly changes. While the change-point problem has been extensively studied for low-dimensional data, advances in data collection technology have produced data sequences of increasing volume and complexity. Motivated by the challenges of modern data, we study a non-parametric framework that utilizes similarity information among observations and can be applied to various data types as long as an informative similarity measure on the sample space can be defined. Analytical p-value approximations are also provided, making the methods easy-off-the-shelf tools for real applications.

Oct. 21, 2020

Model-free Feature Screening and FDR Control with Knockoff Features

Yuan Ke : 4 p.m. in Zoom
Abstract We proposes a model-free and data-adaptive feature screening method for ultra-high dimensional data. The proposed method is based on the projection correlation which measures the dependence between two random vectors. This projection correlation based method does not require specifying a regression model, and applies to data in the presence of heavy tails and multivariate responses. It enjoys both sure screening and rank consistency properties under weak assumptions. A two-step approach, with the help of knockoff features, is advocated to specify the threshold for feature screening such that the false discovery rate (FDR) is controlled under a pre-specified level. The proposed two-step approach enjoys both sure screening and FDR control simultaneously if the pre-specified FDR level is greater or equal to 1/s, where s is the number of active features. The superior empirical performance of the proposed method is illustrated by simulation examples and real data applications.

Oct. 28, 2020

Truncated latent gaussian copula model for zero-inflated data

Irina Gaynanova : 4 p.m. in Zoom
Abstract A great number of multivariate statistical methods, such as principal component analysis, discriminant analysis, canonical correlation analysis and graphical lasso to name a few, require the estimate of covariance or correlation matrix of variables as one of the inputs. It is typical to use Pearson sample correlation matrix, which works well at capturing dependencies between normally distributed variables. In this work we consider the problem of estimating dependencies between zero-inflated measurements, which arise in miRNA data, microbiome data, physical activity data, etc. We propose truncated latent Gaussian copula to model the data with excess zeroes, which allows us to derive a rank-based estimator of latent correlation matrix without the estimation of marginal transformation functions. The new methodology is applied for the analysis of associations between gene expression and microRNA data of breast cancer patients, and for inferring the conditional independence graph in quantitate gut microbiome data.

Nov. 4, 2020

Functional sufficient dimensions reduction via weak conditional moments

Jun Song : 4 p.m. in Zoom
Abstract In this talk, a general theory and estimation methods for functional linear sufficient dimension reduction will be presented, where both the predictor and the response can be random functions or even vectors of functions. Unlike the existing dimension reduction methods, our approach does not rely on the estimation of conditional mean and conditional variance. Instead, it is based on a new statistical construction --the weak conditional expectation, which is based on Carleman operators and their inducing functions. Weak conditional expectation is a generalization of conditional expectation. Its key advantage is to replace the projection on to an L2-space -- which defines conditional expectation -- by projection on to an arbitrary Hilbert space, while still maintaining the unbiasedness of the related dimension reduction methods. This flexibility is particularly important for functional data, because attempting to estimate a full-fledged conditional mean or conditional variance by slicing or smoothing over the space of vector-valued functions may be inefficient due to the curse of dimensionality. We evaluated the performances of our new methods by simulation and in several applied settings.

Nov. 11, 2020

RaSE: Random Subspace Ensemble Classification

Yang Feng : 4 p.m. in Zoom
Abstract We propose a new model-free ensemble classification framework, Random Subspace Ensemble (RaSE), for sparse classification. In the RaSE algorithm, we aggregate many weak learners, where each weak learner is a base classifier trained in a subspace optimally selected from a collection of random subspaces. To conduct subspace selection, we propose a new criterion, ratio information criterion (RIC), based on weighted Kullback-Leibler divergences. The theoretical analysis includes the risk and Monte-Carlo variance of RaSE classifier, establishing the weak consistency of RIC, and providing an upper bound for the misclassification rate of RaSE classifier. An array of simulations under various models and real-data applications demonstrate the effectiveness of the RaSE classifier in terms of low misclassification rate and accurate feature ranking. The RaSE algorithm is implemented in the R package RaSEn on CRAN. This is joint work with Ye Tian.

Nov. 18, 2020

Huber Regression and its Degrees of Freedom

Peng Zeng : 4 p.m. in Zoom
Abstract Huber regression utilizes the Huber loss instead of the common squared loss to achieve the robustness against outliers. It can be regarded as somewhere in the middle of least squares estimate and least absolute deviation. In this talk, we discuss a family of regularized Huber regression models for simultaneous model fitting and variable selection. The prior domain knowledge can be incorporated as linear constraints on parameters. The number of degrees of freedom is a measure of the effective number of parameters used to fit a regression model. It has been used in information criteria for model selection. We derive a formula for the number of degrees of freedom for regularized Huber regression with linear constraints. Simulation studies and real examples are used to demonstrate the application and performance of the proposed methods.

Nov. 25, 2020

Additive Regression for Non-Euclidean Data

Jeong Min Jeon : 4 p.m. in Zoom
Abstract Analyzing non-Euclidean data is becoming an important topic in modern statistics, as various non-Euclidean data are emerging. However, it is not transparent how one can analyze such non-Euclidean data in many subject areas. In this talk, we introduce a general regression method for analyzing many types of non-Euclidean data. In particular, we consider additive models with some metric-space-valued predictors and Hilbertian responses. The predictors in our setting cover any finite-dimensional-Hilbert-space-valued predictors and Riemannian-manifold-valued predictors. Hence, they allow for Euclidean, compositional, circular, spherical and shape-valued predictors. The response setting is broad as well covering Euclidean, compositional, functional and density-valued responses. We present several real data analysis which show the wide applications of our method. We also present its asymptotic theory.

Dec. 2, 2020

Functional Regression with Mixed Predictors

Daren Wang : 4 p.m. in Zoom
Abstract We consider a general functional regression model, allowing for both functional and high-dimensional vector predictors. Based on this general setting, we propose a penalized least squares estimator in reproducing kernel Hilbert spaces (RKHS), where the penalties enforce both smoothness and sparsity on the functional estimator. We also show that the excess prediction risk of our estimator is minimax optimal under this general model setting. Our analysis reveals an interesting phase transition phenomenon and the optimal excess risk is determined jointly by the sparsity and the smoothness of the functional regression coefficients.

Jan. 13, 2021

Statistical Learning for High-dimensional Tensor Data

Anru Zhang : 4 p.m. in Zoom
Abstract The analysis of tensor data has become an active research topic in this area of big data. Datasets in the form of tensors, or high-order matrices, arise from a wide range of applications, such as financial econometrics, genomics, and material science. In addition, tensor methods provide unique perspectives and solutions to many high-dimensional problems, such as topic modeling and high-order interaction pursuit, where the observations are not necessarily tensors. High-dimensional tensor problems generally possess distinct characteristics that pose unprecedented challenges to the data science community. There is a clear need to develop new methods, efficient algorithms, and fundamental theory to analyze the high-dimensional tensor data. In this talk, we discuss some recent advances in high-dimensional tensor data analysis through the consideration of several fundamental and interrelated problems, including tensor SVD and tensor regression. We illustrate how we develop new statistically optimal methods and computationally efficient algorithms that exploit useful information from high-dimensional tensor data based on the modern theories of computation, high-dimensional statistics, and non-convex optimization. Through tensor SVD, we are able to achieve good performance in the denoising of 4D scanning transmission electron microscopy images. Using tensor regression, we are able to use MRI images for the prediction of attention-deficit/hyperactivity disorder.

Jan. 27, 2021

Practices in Statistical Analysis in Psychology Studies and Mental Health Research

Dr. Tao Liu : 4 p.m. in Zoom
Abstract In this seminar, the speaker will present the current practices and trends in data analysis in psychology studies. Research on psychology research methods have found that recent empirical studies published in psychology journals are employing more varied and advanced statistical techniques than were employed previously. The most prevalent statistical analysis methods will be presented, along with the trend of data analysis methods in the past few decades. The presenter will use clinical mental health studies to illustrate such changes, including recent studies during COVID-19 pandemic that investigated the impacts of the public health crisis on psychological wellbeing and social attitudes. Presenter will also provide information of public resources for mental health services and self-care.

Feb. 3, 2021

Efficient Batch Policy Learning in Markov Decision Processes

Zhengling Qi : 4 p.m. in Zoom
Abstract In this talk, I will discuss the batch (off-line) reinforcement learning problem in infinite horizon Markov Decision Processes. Motivated by mobile health applications, we focus on learning a policy that maximizes the long-term average reward. Given limited pre-collected data, we propose a doubly robust estimator for the average reward and show that it achieves statistical efficiency bound. We then develop an optimization algorithm to compute the optimal policy in a parametrized stochastic policy class. The performance of the estimated policy is measured by the difference between the optimal average reward in the policy class and the average reward of the estimated policy. Under some technical conditions, we establish a strong finite-sample regret guarantee in terms of total decision points, demonstrating that our proposed method can efficiently break the curse of horizon. Finally, the performance of the proposed method is illustrated by simulation studies.

Feb. 10, 2021

Pattern Detection for High-Frequency Financial Time Series via Clustering and Bi-Clustering

Jian Zou : 4 p.m. in Zoom
Abstract Exploring high frequency transaction level financial data is of considerable interest to researchers and investors. The extra amount of information contained in high-frequency data and keen interests in high-frequency finance motivate researchers to study dynamic patterns of comovement over multiple trading days. In this paper, we have developed a series of clustering and biclustering algorithms based on mutual information for high frequency financial time series. We examine the co-movement probabilities of selected m-tuples of stocks over multiple trading days under different metrics. Additionally, we propose a unified framework to describe patterns and monitor the structure of high-dimensional daily or weekly time series that track linkages between any given m-tuple of stocks over a long time period.

Feb. 17, 2021

Functional Models for Time Varying Random Objects

Paromita Dubey : 4 p.m. in Zoom
Abstract In recent years, samples of time-varying object data such as time-varying networks that are not in a vector space have been increasingly collected. These data can be viewed as elements of a general metric space that lacks local or global linear structure and therefore common approaches that have been used with great success for the analysis of functional data, such as functional principal component analysis, cannot be applied directly. In this talk, I will propose some recent advances along this direction. First, I will discuss ways to obtain dominant modes of variations in time varying object data. I will describe metric covariance, a novel association measure for paired object data lying in a metric space (\Omega d) that we use to define a metric auto-covariance function for a sample of random \Omega -valued curves, where \Omega generally will not have a vector space or manifold structure. The proposed metric auto-covariance function is non-negative definite when the squared metric d^2 is of negative type. Then the eigenfunctions of the linear operator with the auto-covariance function as kernel can be used as building blocks for an object functional principal component analysis for \Omega-valued functional data, including time-varying probability distributions, covariance matrices and time-dynamic networks. Then I will describe how to obtain analogues of functional principal components for time-varying objects by applying Fréchet means and projections of distance functions of the random object trajectories in the directions of the eigenfunctions, leading to real-valued Fréchet scores and object valued Fréchet integrals. This talk is based on joint work with Hans-Georg Müller.

Feb. 24, 2021

Statistical Modeling and Inference for Next-Generation Functional Data

Guannan Wang : 4 p.m. in Zoom
Abstract With the rapid growth of modern technology, many large-scale imaging studies have been or are being conducted to collect massive datasets with large volumes of imaging data, thus boosting the investigation of "next-generation functional data." These enormous collections of imaging data contain interesting information and valuable knowledge, whichhas raised the demand for further advancement in functional data analysis. In this talk, we mainly focus on modeling and inference of the next-generation functional data. We propose using flexible multivariate splines over triangulation or tetrahedral partitions to handle irregular domain of the images that are common in brain imaging studies and in other biomedical imaging applications. The proposed spline estimators are shown to be consistent and asymptotically normal under some regularity conditions. We also provide a computationally efficient estimator of the covariance function and derive its uniform consistency. Finally, we discuss the inferential capabilities of the proposed method. To be more specific, we develop simultaneous confidence corridors for the mean of the next-generation functional data. The procedure is also extended to the two-sample case in which we focus on comparing the mean functions of random samples drawn from two populations. The proposed method is applied to analyze brain Positron Emission Tomography (PET) data of Alzheimer's Disease.

March 3, 2021

Prediction with Spatially Dependent Functional Covariates

Yeonjoo Park : 4 p.m. in Zoom
Abstract We present a novel spatial model that predicts scalar responses based on functional predictors observed at spatial locations. We incorporate two spatial components in the modeling, (i) spatial correlation between infinite-dimensional functional predictors and (ii) spatially heterogeneous associations between responses and functional covariates at different locations, by introducing a spatially varying functional coefficient model. It allows the functional coefficients to vary with location. To preserve spatial continuity on the low dimensional representation of functional predictors, we employ nonparametric data-adaptive functions for basis expansion under a Bayesian framework and place spatial priors on projection coefficients. We further propose the spatial variable selection, which allows spatially heterogeneous sets of non-null coefficients over locations by borrowing information across neighbors. The basis function estimation, model parameter estimation, and model selection can be jointly performed through Bayesian hierarchical modeling. For the prediction on new observations, we propose the unified approach which enables the estimation of nonparametric basis functions adaptive to new functional predictors and simultaneously draws predictive values from posterior prediction distribution in MCMC implementation. The model performance is demonstrated in simulation studies and an application to a crop yield prediction.

March 10, 2021

On sufficient graphical models

Bing Li : 4 p.m. in Zoom
Abstract We introduce a Sufficient Graphical Model by applying the recently developed nonlinear sufficient dimension reduction techniques to the evaluation of conditional independence. The graphical model is nonparametric in nature, as it does not make distributional assumptions such as the Gaussian or copula Gaussian assumptions. However, unlike a fully nonparametric graphical model, which relies on the high-dimensional kernel to characterize conditional independence, our graphical model is based on conditional independence given a set of sufficient predictors with a substantially reduced dimension. In this way we avoid the curse of dimensionality that comes with a high-dimensional kernel. We develop the population-level properties, convergence rate, and variable selection consistency of our estimate. By simulation comparisons and an analysis of the DREAM 4 Challenge data set, we demonstrate that our method outperforms the existing methods when the Gaussian or copula Gaussian assumptions are violated, and its performance remains excellent in the high-dimensional setting.

March 17, 2021

Multiple imputation methods for unknown stage at diagnosis in cancer data

Pradeep Singh : 4 p.m. in Zoom
Abstract The National Cancer Institute and most states keep a cancer data registry so that it can be used by researchers and policy makers to make better healthcare decisions. This data can have missing observations for one or more variables. In particular, the correct stage at diagnosis is sometimes missing from the data due to various reasons. To use the data, different strategies have been used. Researchers often delete individuals from the study who had missing values from even one variable. Another method is to impute the missing values. There are several methods proposed to impute missing values of quantitative variables. But for categorical variables, there have been few methods proposed. Van der Palm, et al. [2016], compared four imputation methods for categorical data. Zhou et al. [2017] has proposed a nonparametric multiple imputation method using the nearest-neighbor approach. This study applied the nonparametric multiple imputation method proposed by Zhou et al. [2017] and a parametric multiple imputation method to lung adenocarcinoma data from the National Cancer Institute. Lung adenocarcinoma is a type of non-small cell lung cancer that typically forms on the outside of the lungs. A Monte Carlo study was done to compare these methods with respect to imputation bias. The study also compared the effect of different levels (10%, 20%, 40%) of missingness, different sizes of the sample, and different fits of the model on these multiple imputation methods.

March 31, 2021

Sparse Modeling of Functional Linear Regression via Fused Lasso with Application to Genotype-by-environment Interaction Studies

Shan Yu : 4 p.m. in Zoom
Abstract The estimator of coefficient functions in an functional linear model (FLM) based on a small number of subjects is often inefficient. To address this challenge, we propose an FLM based on fused learning. This talk will describe a sparse multi-group FLM to simultaneously estimate multiple coefficient functions and identify groups, such that coefficient functions are identical within groups and distinct across groups. By borrowing information from relevant subgroups of subjects, our method enhances estimation efficiency while preserving heterogeneity in model parameters and coefficient functions. We use an adaptive fused lasso penalty to shrink coefficient estimates to a common value within each group. To enhance computation efficiency and incorporate neighborhood information, we propose to use graph-constrained adaptive lasso with a highly efficient algorithm. This talk will use two real data examples to illustrate the applications of the proposed method on genotype-by-environment interaction studies. This talk features joint work with Aaron Kusmec, Lily Wang, and Dan Nettleton.

April 7, 2021

Single index models with regularized matrix coefficients

Luo Xiao : 4 p.m. in Zoom
Abstract Single index models extend standard linear models to account for non-linearity between multivariate predictors and responses. We study single index models where the unknown coefficients can be formulated as a matrix and enforce regularization term(s) on the coefficient matrix to induce meaningful structure, e.g., sparsity and low-rank. We propose an iterative estimation procedure in which an alternating direction method of multipliers (ADMM) algorithm is employed to accommodate multiple regularization terms. We focus on two particular models: scalar response on matrix predictor model and multivariate response on multivariate predictor model. We apply the former model to study nonlinear association between functional connectivity networks and fluid intelligence, and the latter model to a genetic association study. The work is based on two papers, "Sparse single index models for multivariate responses” which is to appear in Journal of Computational and Graphical Statistics and “Single index models with functional connectivity network predictors”, which has been tentatively accepted by Biostatistics.

April 14, 2021

Locally Weighted Nearest Neighbor Classifier and Its Theoretical Properties

Guan Yu : 4 p.m. in Zoom
Abstract Weighted nearest neighbor (WNN) classifiers are fundamental non-parametric classifiers for classification. They have become the methods of choice in many applications where limited knowledge of the data generation process is available a priori. There exists a vast room of flexibility in the choice of weights for the neighbors in a WNN classifier. In this talk, I will introduce a new locally weighted nearest neighbor (LWNN) classifier, which adaptively assigns weights for different test data points. Given a training data set and a test data point x0, the weights for classifying x0 in LWNN is obtained by minimizing an upper bound of the conditional expected estimation error of the regression function at x0. The resultant weights have a neat closed-form expression, and therefore the computation of LWNN is more efficient than some existing adaptive WNN classifiers that require estimating the marginal feature density. Like most other WNN classifiers, LWNN assigns larger weights for closer neighbors. However, in addition to the ranks of neighbors' distances, the weights in LWNN also depend on the raw values of the distances. Our theoretical study shows that LWNN achieves the minimax rate of convergence of the excess risk, when the marginal feature density is bounded away from zero. In the general case with an additional tail assumption on the marginal feature density, the upper bound of the excess risk of LWNN matches the minimax lower bound up to a logarithmic term.

April 21, 2021

Brain Regions Identified as Being Associated With Verbal Reasoning Through the Use of Imaging Regression via Internal Variation

Xuan Bi : 4 p.m. in Zoom
Abstract Brain-imaging data have been increasingly used to understand intellectual disabilities. Despite significant progress in biomedical research, the mechanisms for most of the intellectual disabilities remain unknown. Finding the underlying neurological mechanisms has proved difficult, especially in children due to the rapid development of their brains. We investigate verbal reasoning, which is a reliable measure of an individual’s general intellectual abilities, and develop a class of high-order imaging regression models to identify brain subregions which might be associated with this specific intellectual ability. A key novelty of our method is to take advantage of spatial brain structures, and specifically the piecewise smooth nature of most imaging coefficients in the form of high-order tensors. Our approach provides an effective and urgently needed method for identifying brain subregions potentially underlying certain intellectual disabilities. The idea behind our approach is a carefully constructed concept called internal variation (IV). The IV employs tensor decomposition and provides a computationally feasible substitution for total variation, which has been considered suitable to deal with similar problems but may not be scalable to high-order tensor regression. Before applying our method to analyze the real data, we conduct comprehensive simulation studies to demonstrate the validity of our method in imaging signal identification. Next, we present our results from the analysis of a dataset based on the Philadelphia Neurodevelopmental Cohort for which we preprocessed the data including reorienting, bias-field correcting, extracting, normalizing, and registering the magnetic resonance images from 978 individuals. Our analysis identified a subregion across the cingulate cortex and the corpus callosum as being associated with individuals’ verbal reasoning ability, which, to the best of our knowledge, is a novel region that has not been reported in the literature. This finding is useful in further investigation of functional mechanisms for verbal reasoning.

April 28, 2021

Change-point detection for COVID-19 time series via self-normalization

Xiaofeng Shao : 4 p.m. in Zoom
Abstract This talk consists of two parts. In the first part, I will review some basic idea of self-normalization (SN) for inference of time series in the context of confidence interval construction and change-point testing in mean. In the second part, I will present a piecewise linear quantile trend model to model infection trajectories of COVID-19 daily new cases. To estimate the change-points in the linear trend, we develop a new segmentation algorithm based on SN test statistics and local scanning. Data analysis for COVID-19 infection trends in many countries demonstrates the usefulness of our new model and segmentation method.

TBA

Xiaofeng Shao : 4 p.m. in Zoom
Abstract TBA

Aug. 25, 2021

Organizational Meeting

: 4 p.m. in Zoom
Abstract We will hold an organizational meeting to welcome everyone in the first week. At this stage, the meeting is going to be a remote format on Zoom. The seminar access link and other information will be sent out through the seminar email list when the date comes closer.

Sept. 8, 2021

Variable Selection for Global Fréchet Regression

Danielle Tucker : 4 p.m. in Zoom
Abstract Global Fréchet regression is an extension of linear regression to cover more general types of responses, such as distributions, networks and manifolds, which are becoming more prevalent. In such models, predictors are Euclidean while responses are metric space valued. Predictor selection is of major relevance for regression modeling in the presence of multiple predictors but has not yet been addressed for Fréchet regression. Due to the metric space valued nature of the responses, Fréchet regression models do not feature model parameters, and this lack of parameters makes it a major challenge to extend existing variable selection methods for linear regression to global Fréchet regression. In this work, we address this challenge and propose a novel variable selection method that overcomes it and has good practical performance. We provide theoretical support and demonstrate that the proposed variable selection method achieves selection consistency. We also explore the finite sample performance of the proposed method with numerical examples and data illustrations.

Sept. 15, 2021

Causal Inference with Continuous Exposures

Ted Westling : 4 p.m. in Zoom
Abstract Much of the literature on estimating causal effects concerns discrete exposures. Recently, there has been increased interest in continuous exposures; that is, exposures that can take an uncountable number of values. Examples of such exposures include air pollution, pre-vaccination antibody responses, and concentrations of harmful chemicals in the blood. In this talk, I will provide an introduction to the area of causal inference with continuous exposures. I will then provide an overview of some of the recent research concerning nonparametric causal inference with continuous exposures, including my own recent and ongoing research. In particular, I will discuss approaches to nonparametric pointwise and global inference on causal dose-response curves, and, time permitting, inference on alternative causal parameters such as the effects of stochastic and incremental interventions.

Sept. 29, 2021

Nonparametric Estimation of Repeated Densities with Heterogeneous Sample Sizes

Xiongtao Dai : 4 p.m. in Zoom
Abstract Functional data analysis concerns a sample of random functions, such as a collection of body growth trajectories. Dimension reduction tools, such as functional principal component analysis, are available to reduce and represent the infinite-dimensional functions. In this work, we are interested in estimating densities as functions, where each density comes from a subpopulation. For example, in the context of epidemiology, the age distributions of patients with different diseases is of central interest, where the disease defines a subpopulation. A key challenge comes from the highly variable sample sizes for different conditions, making the estimation of age profiles difficult for rare conditions. We propose a fully data-driven approach to estimate the densities without the need of specifying the parametric form of the density families. The idea is to map the density functions to a Hilbert space and then apply functional data analytic methods so as to derive low-dimensional approximates. I will show that the proposed methods yield interpretable results and are efficient for modeling electronic medical records and extreme rainfall.

Oct. 6, 2021

Regularized Low-Rank Matrix Regression

Hsin-Hsiung Huang : 4 p.m. in Zoom
Abstract While matrix variate regression models have been studied in many existing works, classical statistical and computational methods for analysis of the regression coefficient estimation are highly affected by ultrahigh dimensional matrix-valued predictors. To address this issue, this paper proposes a framework of matrix variate regression methods, based on a rank-constraint optimization problem and its alternating gradient descent algorithm. In particular, we consider three low-rank matrix variate regression models including ordinary matrix regression, robust matrix regression, and matrix logistic regression, and we establish the convergence property and statistical consistency of the proposed estimator under these three models. The rank constraint effectively reduces the number of parameters in the model, and as a result, compared with existing methods based on regularization, our method has a better theoretical consistency rate. The experimental results show that the proposed algorithms are effective and efficient under various settings.

Oct. 13, 2021

Towards practical estimation of Brenier maps

Jonathan Niles-Weed : 4 p.m. in Zoom
Abstract Given two probability distributions in R^d, a transport map is a function which maps samples from one distribution into samples from the other. For absolutely continuous measures, Brenier proved a remarkable theorem identifying a unique canonical transport map, which is monotone in a suitable sense. We study the question of whether this map can be efficiently estimated from samples. The minimax rates for this problem were recently established by Hutter and Rigollet (2021), but the estimator they propose is computationally infeasible in dimensions greater than three. We propose two new estimators---one minimax optimal, one not---which are significantly more practical to compute and implement. The analysis of these estimators is based on new stability results for the optimal transport problem and its regularized variants. Based on joint work with Manole, Balakrishnan, & Wasserman and with Pooladian.

Oct. 20, 2021

Information-based Optimal Subdata Selection for Clusterwise Linear Regression Model

Yanxi Liu : 4 p.m. in Zoom
Abstract As the data size increases rapidly, the relationship between input and output variables may not be homogeneous anymore. Conventional statistical models such as generalized linear models (GLMs) may not be well-suited to heterogeneous relationships. Using a Mixture of Expert models is a good solution. The Mixture of Expert models can combine different statistical models to detect heterogeneous patterns while maintaining the benefits of conventional statistical modeling techniques. However, it needs a considerable amount of computer resources, particularly when working with big data. To address this issue, an attractive idea is to analyze a subsample of the data retaining the rich information of the full data. Information-Based Optimal Subdata Strategy (IBOSS), proposed by Wang et al. (2019), is such a strategy. The IBOSS strategy captures most of the relevant information in the full data through a judicious selection of the subdata by "maximizing" the Fisher information matrix. This project aims to develop an algorithm for the Clusterwise Linear Regression model, a type of Mixture of Experts, to select subdata based on IBOSS strategy. However, the Fisher information matrix of the model has no explicit form, which is a major challenge of the work. To overcome this challenge, we propose a surrogate matrix which is proved to be asymptotically equivalent to the Fisher information matrix, and it is used to construct the IBOSS subdata. Further, the proposed subdata selection is proved to be asymptotically optimal, i.e., no other method is statistically more efficient than the proposed one when the full data size is large.

Oct. 27, 2021

Evidence factors from multiple, possibly invalid, instrumental variables

Youjin Lee : 4 p.m. in Zoom
Abstract Instrumental variables have been widely used to estimate the causal effect of a treatment on an outcome in the presence of unmeasured confounders. When several instrumental variables are available and the instruments are subject to possible biases that do not completely overlap, a careful analysis based on these several instruments can produce orthogonal pieces of evidence (i.e., evidence factors) that would strengthen causal conclusions when combined. We develop several strategies, including stratification, to construct evidence factors from multiple candidate instrumental variables when invalid instruments may be present. Our proposed methods deliver nearly independent inferential results each from candidate instruments under the more liberally defined exclusion restriction than the previously proposed reinforced design. We apply our stratification method to evaluate the causal effect of malaria on stunting among children in Western Kenya using three nested instruments that are converted from a single ordinal variable. Our proposed stratification method is particularly useful when we have an ordinal instrument of which validity depends on different values of the instrument. This is based on joint work with Anqi Zhao, Dylan Small, and Bikram Karmarkar.

Nov. 3, 2021

Balancing Inferential Integrity and Disclosure Risk via Model Targeted Masking and Multiple Imputation

Bei Jiang : 4 p.m. in Zoom
Abstract There is a growing expectation that data collected by government-funded studies should be openly available to ensure research reproducibility, which also increases concerns about data privacy. A strategy to protect individuals' identity is to release multiply imputed (MI) synthetic datasets with masked sensitivity values (Rubin, 1993). However, information loss or incorrectly specified imputation models can weaken or invalidate the inferences obtained from the MI-datasets. We propose a new masking framework with a data-augmentation (DA) component and a tuning mechanism that balances protecting identity disclosure against preserving data utility. Applying it to a restricted-use Canadian Scleroderma Research Group (CSRG) dataset, we found that this DA-MI strategy achieved a 0% identity disclosure risk and preserved all inferential conclusions. It yielded 95% confidence intervals (CIs) that had overlaps of 98.5% (95.5%) on average with the CIs constructed using the full, unmasked CSRG dataset in a work-disability (interstitial lung disease) study. The CI-overlaps were lower for several other methods considered, ranging from 73.9% to 91.9% on average with the lowest value being 28.1%; such low CI-overlaps further led to some incorrect inferential conclusions. These findings indicate that the DA-MI masking framework facilitates sharing of useful research data while protecting participants' identities. This a joint work with Adrian Raftery (University of Washington), Russel Steele (McGill University) and Naisyin Wang (University of Michigan).

Nov. 10, 2021

Minimax Nonparametric Multi-sample Test Under Smoothing

Pang Du : 4 p.m. in Zoom
Abstract We consider the problem of comparing probability densities between two groups. A new probabilistic tensor product smoothing spline framework is developed to model the joint density of two variables. Under such a framework, the probability density comparison is equivalent to testing the presence/absence of interactions. We propose a penalized likelihood ratio test for such interaction testing and show that the test statistic is asymptotically chi-square distributed under the null hypothesis. Furthermore, we derive a sharp minimax testing rate based on the Bernstein width for nonparametric two-sample tests and show that our proposed test statistics is minimax optimal. In addition, a data-adaptive tuning criterion is developed to choose the penalty parameter. Simulations and real applications demonstrate that the proposed test outperforms the conventional approaches under various scenarios.

Nov. 17, 2021

Subspace learning for high dimensional tensor data

Yuefeng Han : 4 p.m. in Zoom
Abstract Motivated by modern scientific research, analysis of tensors (multi-dimensional arrays) has emerged as one of the most important and active areas in modern statistics and data science. High-dimensional tensor data routinely arise in a wide range of applications, such as economics, genetics, microbiome studies, brain imaging, and hyperspectral imaging, due to modern data collection capabilities. In many of these settings, the observed tensors are of high dimension and high order, but the important information may lie in dimension-reduced subspaces induced by various structural conditions. This talk aims to develop new methodologies and theories from a perspective of subspace learning. The talk is divided into two parts. In the first part, we introduce a factor approach for analyzing high dimensional dynamic tensors, in a form similar to Tucker tensor decomposition. We propose two estimation methods that are based on the tensor unfolding of lagged cross-product and iterative orthogonal projections of the original dynamic tensors. We also establish computational and statistical guarantees of the proposed methods. In the second part, we investigate a tensor factor model with a CP type low-rank tensor structure. We develop a new computationally efficient estimation procedure, which includes a warm-start initialization and an iterative concurrent orthogonalization scheme. We show that the iterative algorithm achieves $\epsilon$-accuracy guarantee within $\log\log(1/\epsilon)$ number of iterations.

Dec. 1, 2021

Minimax Off-Policy Evaluation for Multi-Armed Bandits

Cong Ma : 4 p.m. in Zoom
Abstract This talk is concerned with the problem of off-policy evaluation in the multi-armed bandit model with bounded rewards. We develop minimax rate-optimal procedures under three different settings. First, when the behavior policy is known, we show that the Switch estimator, a method that alternates between the plug-in and importance sampling estimators, is minimax rate-optimal for all sample sizes. Second, when the behavior policy is unknown, we analyze performance in terms of the competitive ratio, thereby revealing a fundamental gap between the settings of known and unknown behavior policies. When the behavior policy is unknown, any estimator must have mean-squared error larger---relative to the oracle estimator equipped with the knowledge of the behavior policy---by a multiplicative factor proportional to the support size of the target policy. Moreover, we demonstrate that the plug-in approach achieves this worst-case competitive ratio up to a logarithmic factor. Third, we initiate the study of the partial knowledge setting in which it is assumed that the minimum probability taken by the behavior policy is known. We show that the plug-in estimator is optimal for relatively large values of the minimum probability, but is sub-optimal when the minimum probability is low. In order to remedy this gap, we propose a new estimator based on approximation by Chebyshev polynomials that provably achieves the optimal estimation error. This is a joint work with Banghua Zhu, Jiantao Jiao and Martin Wainwright.

Feb. 2, 2022

An organizing meeting

TBA : 4 p.m. in Zoom

Feb. 16, 2022

Logarithmic Sobolev Inequalities on Non-isotropic Heisenberg Groups

Liangbing Luo : 4 p.m. in Zoom
Abstract In this talk, I will discuss logarithmic Sobolev inequalities with respect to a heat kernel measure on finite-dimensional and infinite-dimensional Heisenberg groups. Such a group is the simplest non-trivial example of a sub-Riemannian manifold. First, I will talk about logarithmic Sobolev inequalities on non-isotropic Heisenberg groups and discuss the dimension (in)dependence of the constants. In this setting, a natural Laplacian is not an elliptic but a hypoelliptic operator. The argument relies on comparing logarithmic Sobolev constants for the three-dimensional non-isotropic and isotropic Heisenberg groups, and tensorization of logarithmic Sobolev inequalities in the sub-Riemannian setting. Moreover, I will mention the application of these results to an infinite-dimensional Heisenberg group.

March 2, 2022

Residual Refitting Inference for High-Dimensional Linear Model

Yumou Qiu : 4 p.m. in Zoom
Abstract We study statistical inference for the effects of multiple covariates of interest simultaneously after adjusting the effects of high-dimensional control variables under a linear model. A residual refitting procedure is proposed which first obtains the residuals from fitting the response variable and the target covariates on the control covariates via regularized estimation, and then refit the residuals from the first step. Hypothesis testing and confidence interval are constructed. The proposed procedure reduces the impact of the potential over-fitting errors from regularized estimation on the inference of the target parameters. It eliminates the prediction errors in the direction of the true regression error, and hence, achieving more accurate size and higher power. Expansions of the proposed statistics are derived without a sparsity condition on the precision matrix of covariates, which show the error reduction property of the residual refitting procedure. Simulation studies and real data analysis for S&P 500 stock returns verify the theoretical results and demonstrate the proposed method has better performance than the existing methods.

March 9, 2022

Project and Forget: Solving Highly Constrained Optimization Problems

Rishi Sonthalia : 4 p.m. in Zoom
Abstract Many important machine learning problems can be formulated as highly constrained convex optimization problems. One important example is metric constrained problems. In this paper, we show that standard optimization techniques can not be used to solve metric constrained problem. To solve such problems, we provide a general active set framework, called Project and Forget, and several variants thereof that use Bregman projections. Project and Forget is a general purpose method that can be used to solve highly constrained convex problems with many (possibly exponentially) constraints. We provide a theoretical analysis of Project and Forge} and prove that our algorithms converge to the global optimal solution and have a linear rate of convergence. In this talk I will go over the main details of the algorithm, the convergence results and applications to metric constrained and non metric constrained problems. For the non metric constrained problem I will present an application to a new formulation of unbalanced optimal transport known as dual regularized optimal transport.

March 30, 2022

Statistical Inference for Functional Linear Quantile Regression

Peijun Sang : 4 p.m. in Zoom
Abstract We propose inferential tools for functional linear quantile regression where the conditional quantile of a scalar response is assumed to be a linear functional of a functional covariate. In contrast to conventional approaches, we employ kernel convolution to smooth the original loss function. The coefficient function is estimated under a reproducing kernel Hilbert space framework. A gradient descent algorithm is designed to minimize the smoothed loss function with a roughness penalty. With the aid of the Banach fixed-point theorem, we show the existence and uniqueness of our proposed estimator as the minimizer of the regularized loss function in an appropriate Hilbert space. Furthermore, we establish the convergence rate as well as the weak convergence of our estimator. As far as we know, this is the first weak convergence result for a functional quantile regression model. Pointwise confidence intervals and a simultaneous confidence band for the true coefficient function are then developed based on these theoretical properties. Numerical studies including both simulations and a data application are conducted to investigate the performance of our estimator and inference tools in finite sample. This is a joint work with my collaborators Zuofeng Shang and Pang Du.

April 6, 2022

Canceled

Roland Molontay : 4 p.m. in Zoom

April 13, 2022

Bayesian Inference via Filtering Equations for Financial Ultra-High Frequency Data

Yong Zeng : 4 p.m. in Zoom
Abstract We propose a general partially-observed framework of Markov processes with marked point process observations for ultrahigh frequency (UHF) transaction price data, allowing other observable economic or market factors. We develop the corresponding Bayesian inference via filtering equations to quantify parameter and model uncertainty. Specifically, we derive filtering equations, which are SPDEs, to characterize the evolution of the statistical foundation such as likelihoods, posteriors, Bayes factors, and posterior model probabilities. Given the computational challenge, we provide a weak convergence theorem, enabling us to employ the Markov chain approximation method to construct consistent, easily-parallelizable, recursive algorithms. The algorithms calculate the fundamental statistical characteristics and are capable of implementing the Bayesian inference in real-time for streaming UHF data via parallel computing for sophisticated models. The general theory is illustrated by specific models built for U.S. Treasury Notes transactions data from GovPX and a Heston stochastic volatility model for stock transactions data. This talk consists of joint works with B. Bundick, G. X. Hu, D. Kuipers, and J. Yin.

April 20, 2022

Efficient and equitable recruitment into clinical trials using the PREDICTEE algorithm: Application to trials of HCV vaccines

Alexander Gutfraind : 4 p.m. in Zoom
Abstract Randomized clinical trials are a pillar of evidence-based medicine, but are often very expensive and recruit a poor representation of the at-risk population. Here we seek to develop a recruitment strategy for trials of vaccines for HCV that would not only decrease the required sample size to achieve adequate statistical power, but also improve the demographic representation of the recruited trial cohort to enhance their equity and generalizability. Using PWID data collected from Chicago, predictive incidence models were trained and applied to a recruitment scheme which aggregates and enrolls candidates in a batchwise manner and incorporates sample size re-estimation. Dynamic weights are applied to generate a numerical score that can be used to assess a candidate’s expected probability of infection and demographic desirability, thus allowing trials to selectively recruit high-incidence PWID who also contribute to the generalizability of the trial. Simulated clinical trial recruitment using this scheme expressed a two- to three-fold increase in HCV incidence among the trial cohort compared to conventional methods. Simultaneously, the demographic composition of the recruited cohort more closely resembled the target population. This recruitment scheme also proved flexible to varying numbers of matched demographic categories, while also being robust to target populations that were highly dissimilar to the recruitment pool. This novel method of trial recruitment presents a promising approach by which costs can be minimized while assuring a high level of demographic representation. This is joint work with Richard Guan Chiu.

April 27, 2022

The Roles of Statisticians at Center of Drug Evaluation and Research, FDA

Jian Zhao : 4 p.m. in Zoom
Abstract This presentation will describe the roles of statisticians at FDA, Center for Drug Evaluation and Research. High level overview of FDA missions and organizations as well as responsibilities of FDA statisticians in the Office of Biostatistics will be presented. IND and NDA review process will be briefly introduced followed by a review case study for illustration. In addition, a variety of opportunities for statisticians working at FDA including public meetings, guidance &amp; policy development, working groups, research, leadership development, will be discussed.

Aug. 31, 2022

An organization meeting

TBA : 4 p.m. in 636 SEO
Abstract We will hold a welcome meeting for everyone. The new students are encouraged to join to get information about the seminar format.

Sept. 7, 2022

A Computational Biologist’s Bayesian Journey

Qunfeng Dong : 4 p.m. in 636 SEO
Abstract Many biomedical researchers including myself lack the formal training in Bayesian statistics, yet we have found the beauty and magic in it. In this talk, I will present three projects: (1) Bayesian modeling to estimate hospitalization risk for COVID-19 patients with comorbidities, (2) a microbiome taxonomic classification method based on Bayes theorem and bootstrapping, and (3) predicting clinical outcomes of metastatic melanoma patients based on the commensal microbiome using a Bayes’ classifier. I will also highlight the limitations of our methods, so that hopefully hardcore mathematicians/statisticians can come up with better solutions. References: 1. Xiang Gao and Qunfeng Dong (2020) A Bayesian Framework for Estimating the Risk Ratio of Hospitalization for People with Comorbidity Infected by the SARS-CoV-2 Virus. Journal of the American Medical Informatics Association, 28 Sept 2020, ocaa246, doi:10.1093/jamia/ocaa246 2. Xiang Gao, Huaiying Lin, Qunfeng Dong (2017); A Dirichlet-Multinomial Bayes Classifier for Disease Diagnosis with Microbial Compositions, mSphere, Volume: 2, Issue: 6. 3. Xiang Gao, Huaiying Lin, Kashi Revanna, Qunfeng Dong (2017) A Bayesian Taxonomic Classification Method for 16S rRNA Gene Sequences with Improved Species-level Accuracy. BMC Bioinformatics 2017 May 10;18(1):247.

Sept. 21, 2022

New 2-parameter families of advanced forecasting functions: seasonal/nonseasonal models, comparison to the exponential smoothing and ARIMA models, and application to stock market data

Nabil Kahouadji : 4 p.m. in 636 SEO
Abstract We introduce twenty-four new two-parameter families of advanced time series forecasting functions, using three forecast estimate methods along with eight optimization criteria. We also introduce the concept of powering and derive non-seasonal and seasonal time series models with examples in education, sales, economics, industry and finance. We compare the performance of our twenty-four functions/models to both exponential smoothing and ARIMA models using non-seasonal and seasonal time series. We show in particular that our models not only do not require a decomposition of a seasonal time series into trend, seasonal and random components, but also leads to substantially lower sum of absolute error and a higher number of closer forecasts than both Holt--Winters and ARIMA models. Finally, we apply and compare the performance of our twenty-four models using five-year stock market data of 467 companies of the S&P500.

Oct. 5, 2022

Perturbed M-Estimation: Ideas from Robust Statistics for Differential Privacy

Roberto Molinari : 4 p.m. in 636 SEO
Abstract Differential privacy (DP) provides an elegant mathematical framework for defining a provable disclosure risk in the presence of arbitrary adversaries: it guarantees that whether an individual is in a database or not, the results of a DP procedure should be similar in terms of their probability distribution. While DP mechanisms are provably effective in protecting privacy, they often negatively impact the utility of the query responses, statistics, and/or analyses that come as outputs from these mechanisms. To address this problem, we use ideas from the area of robust statistics, which aims at reducing the influence of outlying observations on statistical inference. Based on the preliminary known links between differential privacy and robust statistics, we modify the objective perturbation mechanism by making use of a new bounded function and define a bounded M-Estimator with adequate statistical properties. The resulting privacy mechanism, named “Perturbed M-Estimation”, shows important potential in terms of improved statistical utility of its outputs as suggested by some preliminary results. These results motivate the current work which is being made in this direction.

Oct. 12, 2022

Potential theory of Dirichlet forms with jump kernels blowing up at the boundary

Renming Song : 4 p.m. in 636 SEO
Abstract In this talk, I will present some recent results on potential theory of Dirichlet forms on the half-space $\mathbb{R}^d_+$ defined by the jump kernel $J(x,y)=|x-y|^{-d-\alpha}{\cal B}(x,y)$, where $\alpha\in (0,2)$ and ${\cal B}(x,y)$ can blow up to infinity at the boundary. The main results include boundary Harnack principle and sharp two-sided Green function estimates. This talk is based on a joint paper with Panki Kim and Zoran Vondracek.

Oct. 19, 2022

Copula-Based Anomaly Scoring of High-Dimensional Data with Application in Telecommunication Networks

Roland Molontay : 4 p.m. in Zoom
Abstract Anomaly detection refers to the process of identifying unexpected objects or patterns, which do not conform to the usual behavior. The detection of “not-normal” observations has attracted a lot of research interest from the machine learning community since it has a wide variety of practical applications. In this talk, I will briefly present an overview of the challenges of unsupervised anomaly detection. I will also present our novel model-based approach that relies on the multivariate probability distribution associated with the observations [1]. Since the rare events are present in the tails of the probability distributions, we use copula functions, which are able to model the fat-tailed distributions well. The presented procedure scales well; it can cope with a large number of high-dimensional samples and also with missing values. I will also demonstrate the usability of the method through a case study, where we analyze a large dataset consisting of the performance counters of a real mobile telecommunication network. [1]: Horváth, G., Kovács, E., Molontay, R., & Nováczki, S. (2020). Copula-based anomaly scoring and localization for large-scale, high-dimensional continuous data. ACM Transactions on Intelligent Systems and Technology (TIST), 11(3), 1-26

Oct. 26, 2022

TBA

Yuehua Cui : 4 p.m. in 636 SEO

Mendelian randomization for causal inference with longitudinal traits

Yuehua Cui : 4 p.m. in Zoom
Abstract Mendelian randomization (MR) uses genetic variants as instrument variables to determine whether an observational association between an exposure and an outcome is causal. The use of Mendelian Randomization reduces regression bias and provides reliable estimate of the likely underlying causal relationship between an exposure and a disease outcome. Most current Mendelian randomization methods are focused on cross-sectional phenotypic traits. Longitudinal studies track the same individual at different time points and have a number of advantages over cross-sectional studies. Motivated by a real study to evaluate the causal effect of hormone level on eating behavior, we propose two MR models to investigate the causal effects in a longitudinal study. In the first model, we assume the current exposure affects the current outcome. In the second model, we assume that the past and/or current exposures contribute to the current outcome. The delayed causal effect is determined by data through a variable selection algorithm. Point-wise and simultaneous testing are developed to assess the existence of causal effects. The method was illustrated via simulation studies and an application to an eating behavior dataset.

Nov. 2, 2022

Improved inference in heteroskedastic regression models with monotone variance function estimation

Soo-Young Kim : 4 p.m. in Zoom
Abstract Various methods for estimating variance functions in heteroscedastic regression models have been developed over the years. We propose methods to estimate a variance function in a heteroskedastic regression model where the variance function is assumed to be smooth and monotone in a predictor variable. The estimation method is based on the maximum likelihood principle, and its computation is carried out through regression splines and the cone projection algorithm. The convergence rate of the estimated variance function is derived, and simulations show that it tends to be closer to the true variance function in a variety of scenarios compared to the existing methods. The estimated variance function from the proposed method provides improved inference about the mean function, in terms of a coverage probability and an average length for an interval estimate. The utility of the method is illustrated through the analysis of real datasets.

Nov. 9, 2022

Bayesian Ultrahigh Dimensional Variable Selection for Mixed-type Multivariate Generalized Linear Models

Hsin-Hsiung Huang : 4 p.m. in Zoom
Abstract Inspired by our recent works on the NSF ATD challenges for spatiotemporal data analysis and modeling and Bayesian clustering research, we investigate whether the Bayesian methods can consistently estimate the model parameters when there are multivariate mixed-type responses. To this end, shrinkage priors are useful for identifying relevant signals in high-dimensional data. We develop a multivariate Bayesian model with shrinkage priors (MBSP) model to mixed-type response generalized linear models (MRGLMs), and we consider a latent multivariate linear regression model associated with the observable mixed-type response vector through its link function. Under our proposed model (MBSP-GLM), multiple responses belonging to the exponential family are simultaneously modeled and mixed-type responses are allowed. We show that the MBSP-GLM model achieves strong posterior consistency when $p$ grows at a subexponential rate with $n$. Furthermore, we quantify the posterior contraction rate at which the posterior shrinks around the true regression coefficients and allow the dimension of the responses $q$ to grow as $n$ grows. This greatly expands the scope of the MBSP model to include response variables of many data types, including binary and count data. To address the non-conjugacy concern, we propose an adaptive sampling algorithm via a P\'{o}lya-gamma data augmentation scheme for the MRGLM estimation. We provide simulation studies and real data examples.

Feb. 22, 2023

Modeling Sparse Discrete Data With Application to Microbiome Data

Hani Aldirawi : 4:30 p.m. in 636 SEO
Abstract Sparse data with a high portion of zeros arise in various disciplines. Modeling sparse data is a challenging and growing research area. In this presentation, we provide statistical methods and tools for analyzing sparse discrete data such as microbiome data. We aim to answer three main questions related to microbiome data. 1) What is the most appropriate probabilistic model for modeling each single microbiome feature? 2) How do we build the most appropriate regression model for each feature when covariates are available? 3) Suppose we have repeated measurements for each subject (longitudinal data), how can we identify the time intervals when the two groups of individuals are significantly different?

March 15, 2023

Supervised Stratified Subsampling: an Approach to Big Data Predictive Analytics

Ming-Chung Chang : 6 p.m. in Zoom
Abstract Predictive analytics encompasses the use of statistical models for prediction. Its power, however, is hindered by the rising amounts of data in recent years. Owing to advanced technology, big data are ubiquitous across disciplines. Such data richness may yield difficulties in predictive analytics either in terms of time cost or numerical stability. In this talk, I will introduce a new subsampling approach to overcome this difficulty for regression problems. The proposed method integrates a nonparametric regression technique and stratified sampling, referred to as supervised stratified subsampling. Theoretical properties are developed to justify this method. Numerical studies show that the proposed method yields good predictions and is against model misspecification.

April 5, 2023

BET and BELIEF

Kai Zhang : 4 p.m. in 636 SEO
Abstract We study the problem of distribution-free dependence detection and modeling through the new framework of binary expansion statistics (BEStat). The binary expansion testing (BET) avoids the problem of non-uniform consistency and improves upon a wide class of commonly used methods (a) by achieving the minimax rate in sample size requirement for reliable power and (b) by providing clear interpretations of global relationships upon rejection of independence. The binary expansion approach also connects the symmetry statistics with the current computing system to facilitate efficient bitwise implementation. Modeling with the binary expansion linear effect (BELIEF) is motivated by the fact that wo linearly uncorrelated binary variables must be also independent. Inferences from BELIEF are easily interpretable because they describe the association of binary variables in the language of linear models, yielding convenient theoretical insight and striking parallels with the Gaussian world. With BELIEF, one may study generalized linear models (GLM) through transparent linear models, providing insight into how modeling is affected by the choice of link. We explore these phenomena and provide a host of related theoretical results. This is joint work with Benjamin Brown and Xiao-Li Meng.

April 12, 2023

CESME: High-Dimensional Clustering via Latent Semiparametric Mixture Models

Boxiang Wang : 4 p.m. in 636 SEO
Abstract Cluster analysis is a fundamental task in machine learning. Several clustering algorithms have been extended to handle high-dimensional data by incorporating a sparsity constraint in the estimation of a mixture of Gaussian models. Though it makes some neat theoretical analysis possible, this type of approach is arguably restrictive for many applications. In this talk, I will introduce a novel latent variable transformation mixture model for clustering in which a mixture of Gaussians is assumed after some unknown monotone data transformation. A new clustering algorithm named CESME is developed for high-dimensional clustering under the assumption that optimal clustering admits a sparsity structure. The use of unspecified transformation makes the model far more flexible than the classical mixture of Gaussians. On the other hand, the transformation also brings quite a few technical challenges to the model estimation as well as the theoretical analysis of CESME. I will present a comprehensive analysis of CESME including identifiability, initialization, algorithmic convergence, and statistical guarantees on clustering. In addition, the convergence analysis has revealed an interesting algorithmic phase transition for CESME, which has also been noted for the EM algorithm in the literature. Leveraging such a transition, a data-adaptive procedure is developed and substantially improves the computational efficiency of CESME. Extensive numerical study and real data analysis show that CESME outperforms the existing high-dimensional clustering algorithms including CHIME, sparse spectral clustering, sparse K-means, sparse convex clustering, and IF-PCA.

April 19, 2023

Local Times of Gaussian Random Fields

Yimin Xiao : 4 p.m. in 636 SEO
Abstract Local times of a Gaussian random field $X = \{X(t),t ∈ \mathbb{R}^N\}$ with values in $\mathbb{R}^d$ carry a lot of analytic and geometric properties about $X$. They also arise naturally in the limit distributions of functionals of integrated and fractionally integrated time series or spatial processes, and in nonlinear cointegrating regression. In this talk, we study the local times of anisotropic Gaussian random fields satisfying strong local nondeterminism with respect to an anisotropic metric. By applying moment estimates for local times, we prove optimal local and global Hölder conditions for the local times for these Gaussian random fields and deduce related sample path properties. These results are closely related to Chung’s law of the iterated logarithm and the modulus of nondifferentiability of the Gaussian random fields. We apply the results to systems of stochastic heat equations with additive Gaussian noise and determine the exact Hausdorff measure function for the level sets of the solution. This talk is based on a joint paper with Davar Khoshnevisan and Cheuk Yin Lee.

April 26, 2023

Bayesian regression of genome-wide association summary statistics

Xiang Zhu : 4 p.m. in 636 SEO
Abstract Large-scale genome-wide association studies (GWAS) have markedly improved our understanding of how common variation in the human genome affects complex traits and diseases. Regression models have been widely used to analyze GWAS, but existing methods often require input data at the individual level, which are hard to obtain due to many administrative issues. Here we provide a Bayesian framework for multiple regression without the need of individual-level data. Specifically, we derive a "Regression with Summary Statistics" (RSS) likelihood function of the multiple regression coefficients based on the univariate regression summary statistics, which are easily available in GWAS. We combine the RSS likelihood with prior distributions that are specifically designed for a wide range of genetic applications, such as heritability estimation, phenotype prediction, pathway enrichment and gene prioritization. To estimate posterior distributions, we develop efficient Markov chain Monte Carlo and variational inference algorithms that scales well with millions of genetic variants. Applying RSS to a host of real-world GWAS summary statistics, we demonstrate that RSS not only achieves similar performance in settings where existing methods work, but also enables many novel analyses and discoveries that existing methods cannot deliver.

Aug. 30, 2023

Organizational Meeting

TBA : 4 p.m. in 636 SEO
Abstract In this welcome meeting, the modality of the seminar will be discussed in person. Everyone is invited. The new students are encouraged to attend to get information and communicate with future colleagues.

Sept. 13, 2023

Cost of Sequential Adaptations and Lower Bound on Mean Squared Error of Post-Adaptation Estimators

Sergey Tarima : 4 p.m. in 636 SEO
Abstract The possibility of early stopping and/or interim sample size re-estimation lead to random sample sizes. When such interim adaptations are informative, the interim decision becomes a component of the sufficient statistic. We decompose the total Fisher Information (FI) into the design FI and a conditional-on-design FI analogous to Molenberghs et al. (2014). We go further, representing the conditional-on-design FI as a weighted linear combination of FIs conditional on realized decisions. This decomposition is useful for quantifying how much mean-squared error will be lost due to planned-informative adaptations. We use The FI unspent by having a planned-informative adaptation to determine the lower bound on mean squared error of post-adaptation estimators [the Cramer-Rao lower bound (1946) and its sequential version suggested by Wolfowitz (1947) are not applicable to such estimators]. Theoretical results are illustrated with simple normal samples collected according to a two-stage design with a possibility of early stopping.

Sept. 27, 2023

Influence of Great Researchers at ISI, Kolkata: From Statistics to Game Theory

T.E.S. Raghavan : 4 p.m. in 636 SEO
Abstract Often graduate students concentrate on accumulating high grades in a variety of graduate courses and years pass by with no real thesis problem in sight. ISI style has always been to let you struggle till you discover your own thesis problem. Inspired by the first chapter of Wald’s seminal work on Statistical decision functions, when I jumped to chapter 2, I found my initial mathematical background was quite inadequate. in steering my initial interest from Wald’s Statistical decision theory into Game theory and to the theory of positive operators, many great researchers at ISI Kolkatta have played a significant role. My talk will walk through some facets of this struggle and the decisive role played by professors C.R Rao, D. Basu, K.R Parthasarathy, V.S Varadarajan and S.R. S Varadhan.

Oct. 18, 2023

Flexible spatio-temporal Hawkes process models for earthquake occurrences

Junhyeon Kwon : 4 p.m. in Zoom
Abstract Hawkes process is one of the most commonly used models for investigating the self-exciting nature of earthquake occurrences. However, seismicity patterns have complicated characteristics due to heterogeneous geology and stresses, for which existing methods with Hawkes process cannot fully capture. This study introduces novel nonparametric Hawkes process models that are flexible in three distinct ways. First, we incorporate the spatial inhomogeneity of the self-excitation earthquake productivity. Second, we consider the anisotropy in aftershock occurrences. Third, we reflect the space–time interactions between aftershocks with a non-separable spatio-temporal triggering structure. For model estimation, we extend the model-independent stochastic declustering (MISD) algorithm and suggest substituting its histogram-based estimators with kernel methods. We demonstrate the utility of the proposed methods by applying them to the seismicity data in regions with active seismic activities.

Nov. 15, 2023

Tail Portfolio

Lingjie Ma : 4 p.m. in 636 SEO
Abstract This paper focuses on portfolio construction at tails. The classical mean-variance portfolio focuses on the first two moments of a return distribution. However, asset returns usually are not Gaussian distributed; rather, they have long and fat tails. As a complementary approach, this paper studies and incorporates tail risk into portfolio optimization. Using 1970 to 2019 S&P 500 data, an empirical study was performed by constructing realistic stock selection investment strategies. The results indicate that quantile optimization produces practical tail portfolios with volatility, diversity and turnover comparable to the classical mean-variance approach. Moreover, a tail portfolio outperforms the mean-variance portfolio consistently over the period studied, particularly when the market is bearish.

Nov. 22, 2023

How to Become an Effective Statistician in a Pharmaceutical Company

Dr. Xianwei Bu : 4 p.m. in 636 SEO
Abstract This oral presentation introduces the clinical trials and new drug development in a pharmaceutical company and a statistician’s roles and responsibilities during the process. Some commonly used statistical analysis methods in clinical trials are described. Priorities and timelines are highlighted as two features for a statistician, followed by a summary of how to become an effective statistician in a pharmaceutical company.

Jan. 31, 2024

Expected Shortfall Regression and Its Applications

Wenxin Zhou : 4 p.m. in 636 SEO
Abstract Expected Shortfall (ES), also known as superquantile or Conditional Value-at-Risk, has been recognized as an important measure in risk analysis and stochastic optimization. In finance, it refers to the conditional expected return of an asset given that the return is below some quantile of its distribution. In this talk, we consider a joint regression framework that simultaneously models the conditional quantile and ES of a response variable given a set of covariates, for which the state-of-the-art approach is based on minimizing a joint loss function that is non-differentiable and non-convex. Motivated by the idea of using Neyman-orthogonal scores to reduce sensitivity with respect to nuisance parameters, we propose statistically robust and computationally efficient two-step procedures for fitting joint quantile and ES regression models under three settings: (i) the classical linear model with $p\ll n$; (ii) high-dimensional sparse models with $p\gg n$, and (iii) nonparametric models with a hierarchical compositional structure. Furthermore, we discuss a more general integrated-quantile regression framework, including ES regression as a special case.

Feb. 14, 2024

The Art of Generative AI

Moontae Lee : 4 p.m. in Zoom
Abstract Large Language Models (LLMs) have transformed Natural Language Processing and the wider spectrum of Artificial Intelligence. Relying on their capability to understand extensive contexts, the groundbreaking innovation lies in tackling multiple tasks by a single formalism: generating contextually coherent and creatively diverse subsequent outputs. This talk overviews recent progress in large language modeling and my journey into the realm of generative AI. Highlighting my focuses on both text and code modalities, the talk delves into my recent research on retrieval-augmented text generation and multilingual code completions. Furthermore, the presentation outlines state-of-the-art ongoing research trajectory on planning, reasoning, aligning, and prospective steps toward self-learning. In conclusion, the talk also touches critical problems in Safety, Ethics, and Governance in Artificial Intelligence.

Feb. 21, 2024

High-dimensional Transformation of Single Transcript Measurements for Identifying Structural Variants in Cancer

Hyo Young Choi : 4 p.m. in Zoom
Abstract Over the last decade, many innovative technologies have generated vast amounts of large-scale biological data. The accumulation of so-called “big data”, especially from next generation sequencing technologies, has created many exciting areas in statistics as well as biology. In particular, statistical tools and machine learning techniques have proven to be critical in cancer genomics, transforming large and complex data into clinically relevant knowledge. While many computational tools have been developed for analyzing such big data, unprecedented challenges remain in turning it into meaningful and actionable insights. This talk primarily concerns the issue of high-dimensional outliers which are often challenging to identify in high-throughput sequencing data due to the special structure of high dimensional space. We introduce a new notion of high dimensional outliers that embraces various types and provides deep insights into understanding the behavior of these outliers based on several asymptotic regimes. As an important application, we introduce a statistical method for unsupervised screening of a range of structural alterations in RNA-seq data. We identify a number of biologically important outliers along with the successful characterization of the subspace associated with outliers, which holds promise for identifying otherwise obscured signals.

March 6, 2024

Causal Inference on Distribution Functions

Dehan Kong : 4 p.m. in 636 SEO
Abstract Understanding causal relationships is one of the most important goals of modern science. So far, the causal inference literature has focused almost exclusively on outcomes coming from the Euclidean space. However, it is increasingly common that complex biomedical datasets are best summarized as data points in non-linear spaces. In this paper, we present a novel framework of causal effects for outcomes from the Wasserstein space of cumulative distribution functions, which in contrast to the Euclidean space, is non-linear. We develop doubly robust estimators and associated asymptotic theory for these causal effects. As an illustration, we use our framework to quantify the causal effect of marriage on physical activity patterns using wearable device data collected through the National Health and Nutrition Examination Survey.

March 13, 2024

Causal inference on biological clocks as personalized biomarkers for Alzheimer’s disease

Yongzhao Shao : 4 p.m. in 636 SEO
Abstract Alzheimer’s disease (AD) stands as the leading cause of dementia and related death. Presently, there is no cure or an effective prevention. AD pathology is highly heterogeneous which may begin 20 years before clinical diagnoses. Effective blood-based biomarkers are desired for early diagnosis and monitoring. Epigenetic clocks and mitotic clocks are individual-level biomarkers of aging and often referred as biological ages. Advanced age is known as the most impactful risk factor for late-onset AD, thus, biological ages have the potential to be the blood-based biomarkers of AD for diagnosis and monitoring in personalized medicine. In this talk, we will discuss the derivation of the epigenetic clocks based on DNA methylation profiles and the accelerations of biological ages as well as the mitotic clocks based on lymphocyte telomere lengths (LTLs). We will discuss causal analyses of the relationships between the epigenetic clocks, mitotic clocks and the risk of AD using Mendelian randomization (MR) analysis. The MR-based analyses using selected variants from genome-wide association studies as instrument variables are largely free of biases due to numerous measured and unmeasured confounding factors. If time permits, we will also discuss some MR-based causal analyses of the utility of heart failure medications in reducing the risk of AD among heart failure survivors. This talk is based on joint research with Dr. Yibeltal Ashebir and Jiehui Xu at NYU Grossman School of Medicine.

March 27, 2024

A Maximin Φp-Efficient Design for Multivariate Generalized Linear Models

Yiou Li : 4 p.m. in 636 SEO
Abstract Experimental designs for a generalized linear model (GLM) often depend on the specification of the model, including the link function, the predictors, and unknown parameters, such as the regression coefficients. To deal with the uncertainties of these model specifications, it is important to construct optimal designs with high efficiency under such uncertainties. Existing methods such as Bayesian experimental designs often use prior distributions of model specifications to incorporate model uncertainties into the design criterion. Alternatively, one can obtain the design by optimizing the worst-case design efficiency with respect to the uncertainties of model specifications. In this work, we propose a new Maximin Φp- Efficient (or Mm-Φp for short) design which aims at maximizing the minimum Φp-efficiency under model uncertainties. Based on the theoretical properties of the proposed criterion, we develop an efficient algorithm with sound convergence properties to construct the Mm-Φp design. The performance of the proposed Mm-Φp design is assessed through several numerical examples.

April 3, 2024

High-dimensional modeling and computation challenges and solutions via Bayesian ultrahigh dimensional variable selection and manifold-constrained optimization

Hsin-Hsiung Huang : 4 p.m. in 636 SEO
Abstract High-dimensional data have become prevalent in all fields that need statistical modeling and data analysis. I introduce my recent research in Bayesian ultrahigh dimensional variable selection, low-rank matrix regression and classification, and robust sufficient dimension reduction (SDR). We develop a Bayesian framework for mixed-type multivariate regression with continuous shrinkage priors that enables joint analysis of mixed continuous and discrete outcomes, allowing variable selection from a large number of covariates (p). We investigate the conditions for posterior contraction, especially when the number of covariates (p) grows exponentially relative to the sample size (n) and develop a two-step approach for variable selection with theorems of a sure screening property and posterior contraction and applications with simulation studies and applications to real datasets. To address challenges in analyzing regression coefficient estimation affected by high-dimensional matrix-valued covariates, we propose a framework for matrix-covariate regression and classification models with a low-rank constraint and additional regularization for structured signals, considering continuous and binary responses, introduce an efficient Riemannian-steepest-descent algorithm for regression coefficient estimation, and prove the consistency of the proposed estimator, showing improvement over existing work in cases where the rank is small with applications through simulations and real datasets of shape images, brain signals, and microscopic leucorrhea images. We propose a novel SDR method robust against outliers using α-distance covariance that effectively estimates the central subspace under mild conditions on predictors without estimating a link function, based on the projection on the Stiefel manifold. We establish convergence properties of the proposed estimation under certain regularity conditions and compare the method's performance with existing SDR methods through simulations and real data analysis, highlighting improved computational efficiency and effectiveness.

April 10, 2024

Cost-effectiveness analysis: A statistical overview

Thomas Mathew : 4 p.m. in 636 SEO
Abstract Identifying treatments or interventions that are cost-effective (more effective at a reasonable cost) is clearly important in health policy decision making, especially in the allocation of health care resources. Various measures of cost-effectiveness that are informative, intuitive and simple to explain have been suggested in the literature. Popular and widely used measures include the incremental cost-effectiveness ratio (ICER), defined as the ratio between the difference of average costs and the difference of average effectiveness in two populations receiving two treatments. The ICER is interpreted as the additional cost per unit of effectiveness gained. Yet another measure is the incremental net benefit (INB), which is the difference between the incremental cost and the incremental effectiveness after multiplying the latter with a "willingness-to-pay" amount. In the talk, I will provide a selected review of the statistical criteria and methodologies for cost-effectiveness analysis. In particular, some of the recently introduced probabilistic criteria will be discussed and examples will be given.

April 17, 2024

Subsampling for Big Data Regression with Measurement Constraints

Lin Wang : 4 p.m. in 636 SEO
Abstract Despite the availability of extensive data sets, it is often impractical to observe the responses or labels for all data points due to various measurement constraints in many applications. To address this challenge, subsampling approaches can be employed to select a subset of design points from a large pool for observation, resulting in substantial savings in labeling costs. In this presentation, I will introduce our recent research on computationally feasible subsampling techniques. Our primary focus is on regression with labeled data, which includes linear regression, ridge regression, and nonparametric additive regression. For these regression tasks, we have developed sampling probabilities that aim to minimize the mean squared error in estimations and predictions. We will demonstrate the effectiveness of our proposed approaches through both theoretical analysis and extensive simulations.

April 24, 2024

Statistics Annual Ceremony

Kyunghee Han : 4 p.m. in 636 SEO
Abstract The Statistical Laboratory will present the 2023-2024 Statistics Graduate Student Service Award, Consulting Award, and Research Award.

Sept. 4, 2024

Organizational meeting

TBA : 4 p.m. in 636 SEO
Abstract In this meeting, the modality of the seminar will be discussed in person. Everyone is invited. The new students are encouraged to attend to get information and communicate with future colleagues.

Sept. 11, 2024

Bayesian Method of Borrowing Study-Level Historical Longitudinal Control Data for Mixed-effects Models with Repeated Measures

Dr. Hong Li : 4 p.m. in 636 SEO
Abstract Bringing historical control information into a new trial appropriately holds the promise of more efficient trial design with more accurate estimates, increased power, and fewer patients allocated to inefficacious control group, provided the historical control data are sufficiently similar to the concurrent control. Interest has been growing over the past few decades in leveraging historical clinical trial on the control arm. However, most of the current historical borrowing methods focus on incorporating patient-level historical control information at only one time point. In this work, we propose a Bayesian hierarchical Mixed effect Models for Repeated Measures (BMMRM) to incorporate aggregated study-level longitudinal historical control estimates into the concurrent trial that collected repeated longitudinal data. The simulation study demonstrates that, as compared to one time point data analysis approach, leveraging longitudinal historical control data produces greater power enhancement and mitigates the power loss when the missing data under missing at random (MAR) mechanism is present. Our work also helps fill the gap of lack of methods borrowing historical longitudinal control data from the published summarized estimates when patient-level control data are not available.

Sept. 18, 2024

Scaling Hawkes processes to one million COVID-19 cases

Seyoon Ko : 4 p.m. in Zoom
Abstract Hawkes stochastic point process models have emerged as valuable statistical tools for analyzing viral contagion. The spatiotemporal Hawkes process characterizes the speeds at which viruses spread within human populations. Unfortunately, likelihood-based inference using these models requires O(N^2) floating-point operations, for N the number of observed cases. Recent work responds to the Hawkes likelihood's computational burden by developing efficient graphics processing unit (GPU)-based routines that enable Bayesian analysis of tens-of-thousands of observations. We build on this work and develop a high-performance computing (HPC) strategy that divides 30 Markov chains between 4 GPU nodes, each of which uses multiple GPUs to accelerate its chain's likelihood computations. We use this framework to apply two spatiotemporal Hawkes models to the analysis of one million COVID-19 cases in the United States between March 2020 and June 2023. In addition to brute-force HPC, we advocate for two simple strategies as scalable alternatives to successful approaches proposed for small data settings. First, we use known county-specific population densities to build a spatially varying triggering kernel in a manner that avoids computationally costly nearest neighbors search. Second, we use a cut-posterior inference routine that accounts for infections' spatial location uncertainty by iteratively sampling latent locations uniformly within their respective counties of occurrence, thereby avoiding full-blown latent variable inference for 1,000,000 infection locations.

Oct. 2, 2024

Definitive Screening Designs: What, Why, & How

Bradley Jones : 4 p.m. in 636 SEO
Abstract Definitive Screening Designs (DSDs) were introduced in 2011. Since then they have become popular for industrial applications. This talk describes what a DSD is. It then explains why engineers prefer them to standard two-level fractional factorial designs. Finally, it shows how to construct them and block them.

Oct. 9, 2024

Weighted shape-constrained estimation with applications to Markov chain autocovariance function estimation

Hyebin Song : 4 p.m. in 636 SEO
Abstract In this talk, I will introduce a novel weighted l2 projection method for estimating covariance functions, with an emphasis on estimation of autocovariance sequences from reversible Markov chains. Shape-constrained estimation of a function with discrete support has been investigated and successfully applied to various application problems. Notably, Berg and Song (2023) connected this idea with uncertainty quantification in Markov chain Monte Carlo (MCMC) samples and proposed a shape-constrained estimator for autocovariance sequences. While the least-squares objective is commonly used in shape-constrained regression, it can be suboptimal due to correlation and unequal variances in the input function. To address this, we introduce a weighted least-squares method that defines a weighted norm on transformed data. Our approach involves transforming input data into the frequency domain and weighting the input sequence based on their asymptotic variances, exploiting the asymptotic independence of periodogram ordinates. I will discuss the computational aspects, theoretical properties, and the improved performance of this method compared to its non-weighted counterpart.

Oct. 16, 2024

Change Point Inference for Non-Euclidean Data Sequences using Distance Profiles

Paromita Dubey : 4 p.m. in 636 SEO
Abstract We introduce a powerful scan statistic and the corresponding test for detecting the presence and pinpointing the location of a change point within the distribution of a data sequence with the data elements residing in a separable metric space (Ω, d). These change points mark abrupt shifts in the distribution of the data sequence as characterized using distance profiles, where the distance profile of an element ω ∈ Ω is the distribution of distances from ω as dictated by the data. This approach is tuning parameter free, fully non-parametric and universally applicable to diverse data types, including distributional and network data, as long as distances between the data objects are available. We obtain an explicit characterization of the asymptotic distribution of the test statistic under the null hypothesis of no change points, rigorous guarantees on the consistency of the test in the presence of change points under fixed and local alternatives and near-optimal convergence of the estimated change point location, all under practicable settings. To compare with state-of-the-art methods we conduct simulations covering multivariate data, bivariate distributional data and sequences of graph Laplacians, and illustrate our method on real data sequences of the U.S. electricity generation compositions and Bluetooth proximity networks.

Oct. 23, 2024

Organizational Effectiveness: A New Strategy to Leverage Multisite Randomized Trials for Valid Assessment

Guanglei Hong : 4 p.m. in 636 SEO
Abstract In education, health, and human services, an intervention program is usually implemented by many local organizations. Determining which organizations are more effective is essential for theoretically characterizing effective practices and for intervening to enhance the capacity of ineffective organizations. In multisite randomized trials, site-specific intention-to-treat (ITT) effects are likely invalid indicators for organizational effectiveness and may lead to inequitable decisions. This is because sites differ in their local ecological conditions including client composition, alternative programs, and community context. Applying the potential outcomes framework, this study proposes a mathematical definition for the relative effectiveness of an organization. The estimand contrasts the performance of a focal organization with those that share the features of its local ecological conditions. The identification relies on relatively weak assumptions by leveraging observed control group outcomes that capture the confounding impacts of alternative programs and community context. We propose a two-step mixed-effects modeling (2SME) procedure. Simulations demonstrate significant improvements when compared with site-specific ITT analyses or analyses that only adjust for between-site differences in the observed baseline participant composition. We illustrate its use through an evaluation of the relative effectiveness of individual Job Corps centers by reanalyzing data from the National Job Corps Study, a multisite randomized trial that included 100 Job Corps centers nationwide serving disadvantaged youths. The new strategy promises to alleviate consequential misclassifications of some of the most effective Job Corps centers as least effective and vice versa.

Nov. 6, 2024

Methods for Informative Censoring in Time-to-Event Data Analysis

Dr. Mandy Jin : 4 p.m. in Zoom
Abstract In oncology clinical trials, subjects prematurely discontinuing from the assigned treatment prior to experiencing an event of interest are often handled by noninformative censoring under censor-at random assumption. Such methods can be challenged with respect to the robustness of the ignorable or noninformative censoring and sensitivity analyses using informative censoring are often required. In a recently published article (Jin and Fang, 2024), reference-based methods (including Jump to Reference and Copy Reference) and tipping point analysis for time-to-event data with possibly informative censoring were proposed. These are novel methods to fit the gap in literature for time-to-event analysis with applications in oncology clinical trials. We will describe and facilitate the implementation of these methods in this presentation. Illustrative examples are provided to demonstrate the reference-based methods and tipping point analysis.

Nov. 13, 2024

Taking Mobile Consumer’s Pulse--An Integrated Analysis of Mobile Application Usage and In-App Advertising Response

Dr. Yingda Lu : 4 p.m. in 636 SEO
Abstract Consumers have increasingly spent more time on mobile applications, and companies have also allocated more resources to advertisement in mobile applications and are actively seeking ways to improve the click-through rate of in-app ads. However, there is a lack of research leveraging consumers’ mobile application usage to understand in-app advertisement. In this study, we develop an integrated model of mobile application usage and in-app advertising response. We use a hidden-Markov model (HMM), which allows consumer involvement in mobile activities to drive temporal changes in both consumer mobile application usage and in-app advertising response. Our framework captures three components that are understudied in previous research on in-app advertising responses: 1) contextual mobile app in which consumers are targeted; 2) long-range correlation in preceding periods and 3) multitasking across mobile apps. To address the challenge of long-range correlation in traditional HMM, we further extend HMM by incorporating a long short-term memory (LSTM) autoencoder into the state transition. Using a unique panel dataset, we find salient temporal patterns and persistence of consumers’ underlying involvement that govern both application usage and advertisement response. Interestingly, consumers’ responses to advertisements follow an inverted-U shape where consumers are most likely to respond to advertisements in a medium state of involvement. Consumers’ advertisement responses are also subject to a contextual effect. For example, consumers are more likely to respond to advertisements when they use Entertainment apps compared with other apps. Our simulation indicates that incorporating mobile usage information, such as a temporal state of involvement and contextual effects at the individual level (viewing history), can significantly improve the effectiveness of targeting strategies. For instance, incorporating contextual effect and multitasking can increase performance by as much as 21.2%. This improvement can be further enhanced with the help of the LSTM autoencoder to address the long-range correlations in HMM. We are the first to connect consumers’ mobile application usage with their in-app ad response.

Nov. 20, 2024

Stage-Aware Learning for Dynamic Treatments

Annie Qu : 4 p.m. in 636 SEO
Abstract Recent advances in dynamic treatment regimes (DTRs) provide powerful optimal treatment searching algorithms, which are tailored to individuals’ specific needs and able to maximize their expected clinical benefits. However, existing algorithms could suffer from insufficient sample size under optimal treatments, especially for chronic diseases involving long stages of decision-making. To address these challenges, we propose a novel individualized learning method which estimates the DTR with a focus on prioritizing alignment between the observed treatment trajectory and the one obtained by the optimal regime across decision stages. By relaxing the restriction that the observed trajectory must be fully aligned with the optimal treatments, our approach substantially improves the sample efficiency and stability of inverse probability weighted based methods. In particular, the proposed learning scheme builds a more general framework which includes the popular outcome weighted learning framework as a special case of ours. Moreover, we introduce the notion of stage importance scores along with an attention mechanism to explicitly account for heterogeneity among decision stages. We establish the theoretical properties of the proposed approach, including the Fisher consistency and finite-sample performance bound. Empirically, we evaluate the proposed method in extensive simulated environments and a real case study for COVID-19 pandemic.

Jan. 15, 2025

BF-BOIN-ET: A backfill Bayesian optimal interval design using efficacy and toxicity outcomes for dose optimization

Kentaro Takeda : 4 p.m. in Zoom
Abstract The primary purpose of a dose-finding trial for novel anticancer agents is to identify an optimal dose (OD), defined as the tolerable dose that has adequate efficacy in unpredictable dose-toxicity and dose-efficacy relationships. The FDA project Optimus reforms the paradigm of dose optimization and recommends that dose-finding trials compare multiple doses to generate these additional data at promising dose levels. The backfill is helpful in settings where the efficacy of a drug does not always increase with the dose level. More information is available at these doses by backfilling patients at lower doses while the trial continues to explore higher doses. This paper proposes a Bayesian optimal interval design using efficacy and toxicity outcomes that allows patients to be backfilled at lower doses during a dose-finding trial while prioritizing the dose-escalation cohort to explore a higher dose. A simulation study shows that the proposed design, the BF-BOIN-ET design, has advantages compared to the other designs in terms of the percentage of correct OD selection, reducing the sample size, and shortening the duration of the trial in various realistic settings.

Feb. 12, 2025

Optimal Exact Design of Experiments: Challenges and Approaches

Radoslav Harman : 4 p.m. in Zoom
Abstract The field of optimal experimental design has traditionally focused on “approximate” designs, which specify a finite set of experimental conditions along with the proportions of trials allocated to each condition. The main advantage of approximate designs is that they allow the use of powerful theoretical and numerical tools from convex optimization. In practical applications, however, “exact” experimental designs are required. These designs determine a finite set of experimental conditions for conducting the trials. Although an exact design can often be derived from an approximate design using rounding algorithms, such procedures typically yield suboptimal results. In this talk, I will first define and review the integer optimization problem underlying optimal exact design for statistical models with uncorrelated observations, emphasizing its theoretical and computational complexity. Next, I will survey various approaches for computing optimal exact designs numerically. In particular, I will highlight popular exchange methods and other heuristic strategies. I will also discuss methods based on mixed-integer mathematical programming formulations, including the latest developments in the field. Finally, I will illustrate these methods on several challenging problems involving exact optimal designs under non-standard experimental constraints.

Feb. 26, 2025

Rare Event Detection by Acquisition-Guided Sampling

Huiling Liao : 4 p.m. in 636 SEO
Abstract Motivated by the challenges in detecting extremely rare failures for sophisticated specifications in circuit design, we consider the problem of detecting regions of interest (ROIs) that consist of specifications with the value of a complex target function for the system performance being below or above a certain pre-specified threshold. Though Bayesian optimization (BO) has been applied to this problem, it is not effective in identifying multiple ROIs as it was originally designed for global optimization and tends to focus on searching the area where the global optimum is most likely to be. In this work, we propose a sampling strategy for fast ROI detection within a limited number of target function evaluations. The sampling distribution is designed so that the probability of a specification being sampled is proportional to the corresponding value of the acquisition function. Such an acquisition-guided sampling algorithm promotes a wider search of the sample space and a simpler incorporation of different criteria to determine the specifications to be evaluated next. To further improve the performance, we propose a new design of the acquisition function and two modifications of existing acquisition functions. Numerical studies on synthetic functions and a real-world circuit design application demonstrate that the proposed method can enjoy a stronger exploration ability provided by sampling and achieve faster ROI detection with higher coverage.

March 19, 2025

Generalised raking and stabilised weights for regression modelling in two-phase samples

Tong Chen : 4 p.m. in Zoom
Abstract In regression models fitted to data from complex survey designs, sampling weights often incorporate non-essential variation, inflating variance estimates. Stabilised weights mitigate this issue by adjusting sampling weights to account for variation explained by covariates. We evaluate the performance of optimal stabilised weights and propose combining the stabilised weights estimator with generalised raking, a class of efficient design-based estimators. This combination improves efficiency by reducing unnecessary weight variation and leveraging information from auxiliary variables. We show this combination can be implemented using the standard statistical package that handles two-phase samples and generalised raking. Simulation studies demonstrate that the proposed estimator enhances precision under realistic two-phase designs, though efficiency gains may be limited in highly informative designs.

April 2, 2025

On the Testing of Statistical Software

Ryan Lekivetz : 4 p.m. in 636 SEO
Abstract Testing statistical software is an extremely difficult task. What is more, for many statistical packages, the developer and test engineer are one and the same, may not have formal training in software testing techniques, and may have limited time for testing. This makes it imperative that the adopted testing approach is both efficient and effective and, at the same time, it should be based on principles that are readily understood by the developer. As it turns out, the construction of test cases can be thought of as a designed experiment (DOE). This talk provides a treatment of DOE principles applied to testing statistical software and includes other considerations that may be less familiar to those developing and testing statistical packages.

April 9, 2025

Defenses Against Backdoor Attacks in Federated Learning and Text Classification

Yao Li : 4 p.m. in 636 SEO
Abstract As machine learning models become increasingly integrated into distributed and language-intensive applications, ensuring their integrity against backdoor attacks is paramount. This talk presents two defense strategies that target vulnerabilities in federated learning and large language models (LLMs). The first part introduces Trusted Aggregation (TAG), a robust defense mechanism for federated learning that leverages a small validation set to estimate permissible updates and filter out malicious contributions. TAG effectively mitigates backdoor risks while preserving task accuracy, even when up to 40% of client updates are adversarial. The second part addresses the threat of syntactic textual backdoor attacks in LLMs. We propose a novel token substitution strategy that alters semantic content while preserving syntactic structures, enabling the detection of both syntax-based and token-based triggers.

April 16, 2025

Integrating Translational Data and Statistical Innovation

Nan Xi : 4 p.m. in 636 SEO
Abstract Combination drug therapies hold significant promise in enhancing treatment efficacy, particularly in fields such as oncology, immunotherapy, and infectious diseases. However, designing clinical trials for these regimens poses unique challenges due to multiple hypothesis testing, shared control groups, and overlapping treatment components that induce complex correlation structures. In this work, we develop a novel statistical framework tailored for early-phase translational combination therapy trials, with a focus on platform trial designs. Our methodology introduces a generalized Dunnett’s procedure that controls false positive rates by accounting for the correlations between treatment arms. Additionally, we propose strategies for power analysis and sample size optimization that leverage preclinical data to estimate effect sizes, synergy parameters, and inter-arm correlations. Simulation studies demonstrate that our approach not only controls various false positive metrics under diverse trial scenarios but also informs optimal allocation ratios to maximize power. A real-data application further illustrates the practical integration of translational preclinical insights into the clinical trial design process. Overall, our framework provides practical and statistically robust guidance for the design of early-phase combination therapy trials, enhancing the efficiency of the bench-to-bedside transition.

April 23, 2025

VIX and VVIX in the Zero-Day-To-Expiration (0DTE) Science Fictional Options Universe

Gilbert W. Bassett : 4 p.m. in 636 SEO
Abstract The VIX--volatility index--is a number. It is transmitted to the world every 15 seconds from downtown Chicago, viewed right outside the window from UIC, at Cboe. It features Dispersion, Probability, Fear, and other topics of interest to statistics students. It is not related to the Black Scholes model, its volatility is not Variance, and the related Volatility of Volatility (VVIX) measures the Fear of Fear. An introduction to the VIX accessible to statistics students is presented via the Science Fictional Options Universe (SFOU) wherein, among other things, Probability does not exist: Never Happened. Not subjective, physical, or frequentist; no dice. People understand that the future is uncertain and have a sense about what is more or less likely, but "likely" in the SFOU is an informal concept, you know what I mean. In our universe it is like Hygge. *Disclosure: I am on the board of directors at the Cboe Futures Exchange that produces the VIX.

Sept. 10, 2025

Organizational meeting

TBA : 4:15 p.m. in 636 SEO

Sept. 24, 2025

Testing composite null hypotheses with high-dimensional dependent data.

Hongyuan Cao : 4:15 p.m. in 636 SEO
Abstract Testing composite null hypotheses is fundamental to many scientific applications, including mediation and replicability analyses, and becomes particularly challenging in high-throughput settings involving tens of thousands of features. Existing high-dimensional composite null hypotheses testing often ignores the dependence structure among features, leading to overly conservative or liberal results. To address this limitation, we develop a four-state hidden Markov model (HMM) for bivariate $p$-value sequences arising from two-study replicability analysis. This model captures local dependence among features and accommodates study-specific heterogeneity. Based on the HMM, we propose a multiple testing procedure that asymptotically controls the false discovery rate (FDR). Extending this framework to more than two studies is computationally intensive, with complexity growing exponentially in the number of studies $n.$ To address this scalability issue, we introduce a novel e-value framework that reduces computational complexity to quadratic in $n,$ while preserving asymptotic FDR control. Extensive simulations demonstrate that our method achieves higher power than existing approaches at comparable FDR levels. When applied to genome-wide association studies (GWAS), the proposed approach identifies novel biological findings that are missed by current methods.

Oct. 1, 2025

Statistical Designs for Network A/B Testing

Qiong Zhang : 4:15 p.m. in Zoom
Abstract A/B testing is an effective method to assess the potential impact of two treatments. For A/B tests conducted by IT companies like Meta and LinkedIn, the test users can be connected and form a social network. Users’ responses may be influenced by their network connections, and the quality of the treatment estimator of an A/B test depends on how the two treatments are allocated across different users in the network. In this talk, I will discuss optimal design criteria based on some commonly used outcome models, under assumptions of network-correlated outcomes or network interference. I will show that the optimal design criteria under these network assumptions depend on several key statistics of the random design vector. I will discuss a framework to develop algorithms that generate rerandomization designs meeting the required conditions of those statistics. I further talk about asymptotic distributions to guide the specification of algorithmic parameters and validate the proposed approach using both synthetic and real-world networks.

Oct. 22, 2025

Bridging Educational Data and Classroom Practice Using Human-Centered AI -- A statistician’s view

Dr. Hongwen Guo : 4:15 p.m. in Zoom
Abstract In the era of digital assessments, large-scale educational data—such as that from NAEP—offers unprecedented opportunities for insight into student learning skills. Yet, the complexity and volume of this data often outpace traditional statistical approaches, calling for a fusion of statistical rigor, data science innovation, and AI-driven modeling. This talk explores a research initiative at ETS, supported by the Gates Foundation, that helps to transform multi-source NAEP data (response, process, and behavioral) into actionable insights for educators. We will discuss how statistics and data science form the foundation for extracting meaningful patterns, visualizing complex data, and how human-centered AI enables scalable, interpretable feedback with subject-matter experts and teachers. The presentation will also reflect on the speaker’s own professional evolution—from classical statistics to data science and to AI applications in the education measurement field —highlighting the synergies between these disciplines. Faculty and graduate students interested in statistical modeling, educational measurement, and AI applications are invited to join the discussion.

Oct. 29, 2025

Subgroup Identification based on Quantitative Objectives for Randomized and Non-Randomized Studies

Yan Sun : 4:15 p.m. in 636 SEO
Abstract Precision medicine is the future of drug development, and subgroup identification plays a critical role in achieving the goal. In this presentation, we propose a powerful end-to-end solution squant (available on CRAN) that explores a sequence of quantitative objectives. The method converts the original study to an artificial 1:1 randomized trial, and features a flexible objective function, a stable signature with good interpretability, and an embedded false discovery rate (FDR) control. We demonstrate its performance through simulation and provide a real data example.

Nov. 5, 2025

Heterogeneous Treatment Effects under Network Interference: A Nonparametric Approach Based on Node Connectivity

Heejong Bong : 4:15 p.m. in 636 SEO
Abstract In network settings, interference between units makes causal inference more challenging as outcomes may depend on the treatments received by others in the network. Typical estimands in network settings focus on treatment effects aggregated across individuals in the population. We propose a framework for estimating node-wise counterfactual means, allowing for more granular insights into the impact of network structure on treatment effect heterogeneity. We develop a doubly robust and non-parametric estimation procedure, KECENI (Kernel Estimator of Causal Effect under Network Interference), which offers consistency and asymptotic normality under network dependence. The utility of this method is demonstrated through an application to microfinance data, revealing the node-wise impact of network characteristics on treatment effects.

Nov. 19, 2025

Regression adjustment with high-dimensional covariates

Dogyoon Song : 4:15 p.m. in Zoom
Abstract Regression adjustment is a classical technique in causal inference that leverages covariates to improve precision of estimators in randomized controlled trials (RCTs) and to adjust for confounding in observational studies. While well-understood in low-dimensional settings, its behavior in modern high-dimensional regimes---where the number of covariates may be comparable to or even exceed the number of observations---remains underexplored. In particular, existing theoretical results are largely asymptotic, often rely on residual-based arguments, and provide limited insights into finite-sample inference especially when $p>n$. In this talk, we revisit regression adjustment for the average treatment effecting (ATE) estimation under complete randomization with many covariates, in a design-based, finite-population framework, via two vignettes. First, we introduce a novel theoretical perspective on the asymptotic properties of regression adjustment through a Neumann-series decomposition, yielding a refined analysis in the $p<n$ regime. Specifically, for ordinary least squares (OLS) regression adjustment, we show that the degree-$d$ Neumann-corrected estimator is asymptotically normal when $p^{d+3}(\log p)^{d+1}=o(n^{d+2})$. This result strictly enlarges the previously reported admissible growth of $p = o(n^{1/2})$ or $p = o(n^{2/3})$ with a single de-biasing step. Second, we present a non-asymptotic analysis of the regression-adjusted ATE estimators that is valid in both $p<n$ and $p>n$ settings. Leveraging concentration of measure tools, we quantify uncertainty without relying on classical asymptotic variance estimation, and further control the design bias of estimators via Stein's method of exchangeable pairs. Time permitting, we will discuss potential extensions and ongoing work.

Dec. 3, 2025

Combining Probability and Non-probability Samples Using Semi-parametric Quantile Regression

Sixia Chen : 4:15 p.m. in Zoom
Abstract Non-probability samples are prevalent in various fields, such as biomedical studies, educational research, and business investigations, owing to the escalating challenges associated with declining response rates and the cost-effectiveness and convenience of utilizing such samples. However, relying on naive estimates derived from non-probability samples, without adequate adjustments, may introduce bias into study outcomes. Addressing this concern, data integration methodologies, which amalgamate information from both probability and non-probability samples, have demonstrated effectiveness in mitigating selection bias. Nonetheless, the efficacy of these methods hinges upon the assumptions underlying the models. This paper introduces innovative and robust data integration approaches, notably a semi-parametric quantile regression-based mass imputation approach and a doubly robust approach that integrates a non- parametric estimator of the participation probability for non-probability samples. Our proposed methodologies exhibit greater robustness compared to existing parametric approaches, particularly concerning model misspecification and outliers. We consider both missing at random and not missing at random scenarios. Theoretical results are established, including variance estimators for our proposed estimators. Through comprehensive simulation studies and real- world applications, our findings demonstrate the promising performance of the proposed estimators in facilitating valid statistical inference. This research contributes to the advancement of robust methodologies for handling non-probability samples, thereby enhancing the reliability and validity of research outcomes across diverse domains.

Feb. 25, 2026

Bayesian non-negative tensor factorization for international trading

Jie Jian : 4:15 p.m. in 636 SEO
Abstract Detecting dependence structures in international trade—such as persistent exporter–importer affinities, and supply-chain clustering—often relies on latent variable models that summarize high-dimensional trading flows. We propose a novel Bayesian non-negative tensor factorization for large, sparse, nonnegative trading tensors with excess zeros and continuous positive measurements. We target settings with millions of entries and extreme sparsity. Each entry follows a spike-and-slab model: a point mass at zero coupled with a gamma–Poisson construction that yields a low-rank nonnegative decomposition via gamma latent factors. The framework provides interpretable mode-specific components and principled uncertainty quantification.

March 11, 2026

Quantile Portfolio Optimization

Lingjie Ma : 4:15 p.m. in 636 SEO
Abstract It is well known that asset returns usually do not follow a normal distribution, rather, they have long and fat tails. This paper focuses on the quantile portfolio methodology, which considers the whole distribution of asset returns and employs expected loss as a risk measurement. In particular, we explore statistical properties of tau risk and propose related theories of quantile portfolio optimization. We also introduce portfolio performance terms for the quantile portfolio framework.

March 18, 2026

Parameter-Expanded Data Augmentation in Probit Models

Dr. Xiao Zhang : 4:15 p.m. in Zoom
Abstract Probit models have been prominent tools to analyze binary/ordinal data, but the computational complexity of maximum likelihood functions presents challenges in their usage. Furthermore, the model identification necessitates the covariance matrix of the latent multivariate normal variables to be a correlation matrix, which brings a rigorous task to develop efficient Markov chain Monte Carlo (MCMC) sampling methods. Data augmentation has been inevitable explored for both identifiable univariate and multivariate probit models. Particularly, it is well-known that parameter-expanded data augmentation (PX-DA) based on non-identifiable models accelerates the convergence and improves the mixing of MCMC components. However, comprehensive investigation has seldom been undertaken, and various algorithms due to incorrectly constructed non-identifiable models further bring obstacles to develop efficient MCMC sampling methods. We tackle this issue by constructing correct non-identifiable models and develop PX-DA algorithms to estimate both univariate and multivariate probit models. Our investigation exhibits that the proposed PX-DA algorithms advance the performance of MCMC sampling considerably and illustrates the essentials of using PX-DA, especially for data with large sample sizes.

April 1, 2026

Low-Rank Distance Covariance for Fréchet Sufficient Dimension Reduction in High-Dimensional Functional Data

Hsin-Hsiung Huang : 4:15 p.m. in 636 SEO
Abstract We develop a new Fréchet sufficient dimension reduction (FdSDR) framework tailored for high-dimensional functional data with complex, metric-space-valued responses. Our main contribution is a low-rank distance covariance criterion that enables scalable, model-free identification of low-dimensional predictor structures while capturing nonlinear dependence. The proposed method is computationally efficient in high dimensions and avoids restrictive distributional assumptions. We establish theoretical guarantees and demonstrate its effectiveness through simulations and real data, providing a practical and flexible approach for modern functional data analysis.