Monday, November 17, 2014

Homework for Session 5 - Batch 3 CBA

Hi all,

Please find here the individual homework for session 5.

Pls ensure you are able to replicate classwork examples with the R code sent before you try this one.

The idea is simple. I will require you to:

  • 1. Pull your facebook (FB) data. Your friends' list. Pls use the Rfacebook package and the instructions from the slides.
  • 2. Run the communities-detection algorithm on it.
  • 3. Paste a screenshot of the network with communities on a slide. Identify the top few clearly identified groups that you can see (like I'd shown for my FB pull in the class slides).
  • 4. Analyze the 5 largest communities you got in terms of (i) size, transitivity, density, centralities, and (ii) meaning (how does the community relate to the ego or focal person).

Submission format and deadline the same as in the past. Save your PPT as (your.full.name).pptx

Any queries etc, contact me.

Thanks.

Sudhir

Saturday, November 15, 2014

Mailbag

Hi all,

Received this in the mail today and responded to it. Am putting up the exchange here coz I think it merits further dissemination.

The email I got:

Hi Sudhir,

I am a CBA Batch -3 student from Section-A

I am facing issues relating to very fundamental meanings of terminology introduced in DCBA. It may be due to the reason that I am not a business guy who is well versed with business terminology.

eg I am not comfortable with the following keywords: construct, dichotomy, Costs of Capital, trickier proposition, business meta-process and so many keywords introduced on Slide 16 Problem Formulations Examples of R.O.s, and then in psychometric scaling

Considering the example of Baskin Robbins. I am not able to get how psychology is coming into the picture here?

I may sound as asking stupid questions but I know I need to do something about it so that I can get comfortable with this subject.

Please provide me some directions..

Thanks,

P

My response:

Hi P,

Let me try to systematically answer what I can.

1. Regarding what a 'construct' means in our context, pls refer to this blogpost (from the PGP class):

http://marketing-yogi.blogspot.in/2014/09/session-2-exposition-what-are.html

2. The definitions of 'dichotomy', 'business processes' and meta-processes, 'cost of capital' can be had from a google search. A dichotomy means a branching into two separate streams. Thus, Data types exhibit a dichotomy - primary versus secondary data etc.

3. 'trickier proposition' is an expression in speech that means "is more problematic" or "is more challenging".

4. Not every construct need have profound psychological drivers. Many are fairly routine and habit driven.

5. The Baskins Robbins example has nothing to do with psychology. Its merely meant to illustrate the primary-secondary dichotomy.

I hope that helps clarify things somewhat, at least. Thanks for reaching out. I might put up this entire exchange on the blog, in case other students are also facing the same problem.

Thanks,

Sudhir

***************************

Updates. Received two more email queries. My responses are also putup below.

Hi Professor,

I am a student of CBA batch 3. I just had a query around the R code for text analysis (filename: textanalysis R code.R), I have gone through the entire code and wanted to understand the last part i.e. Bayes Factor Model selection and thereon. Can you kindly guide me on this?

I am not able to conceptually grasp the concept of Factor Model and the output from that point onwards.

Look forward to hearing from you.

RT

My response:

Hi RT,

>> I have gone through the entire code and wanted to understand the last part i.e. Bayes Factor Model selection and thereon. Can you kindly guide me on this?

Your query concerns what we call 'model selection' in statistics. A model is a set of relations which we fit upon data to explain them and/or make predictions about them. However, there maybe multiple models that fit the same data.

One way to sort through this multiplicity of models and select the *best* one is to first find how well each model 'fits' the data (i.e. has the least squared error). Accordingly, various 'goodness of fit' criteria have been developed and deployed. The Bayes Factor is one such, very important fit metric in Bayesian statistics.

For our purposes, just take the model results and use them to select the model with the optimal number of components (optimal, as decided by the log bayes factor). Going beyond that would be beyond the scope of the DC course. Wikipedia and other web resources are available however, in case you want to do a deep-dive into fit statistics in general and Bayes Factors in particular.

>> I am not able to conceptually grasp the concept of Factor Model and the output from that point onwards.

When we 'factorize' something (say, a), we break it down into pieces (say, b,c and d) such that the product of b*c*d will yield a back.

In general, any number can be 'factorized' into a product of primes. Similarly, when we factorize a matrix, we break it down into 'factors' whose product yields the original matrix again.

We took the TDM and 'factorized' it (conceptually only the LDA is more complex in its assumptions and its estimation) into 'factors' - terms that together can be interpreted as topics.

For our purposes, all we need to know is that using the latent topic factor model, we 'broke down' the corpus into distinct 'topics' or themes that can be interpreted and used for further analysis.

I hope that helps clarify.

Sudhir

Another one below:

Respected Professor,

I'm a CBA student from technology background. I need your help regarding data collection:

1. Is there any book that I can refer? I feel I'm lost with so much of info/topics. Also with no audio for first class, it seems I don't have way to revisit the fundamentals discussed.

2. It will be extremely helpful if you can please provide some practice papers and solutions. (hope that's possible)

3. Could you please also clarify whether Facbook assignment is group or individual H.W.?

SP

My response:

Hi SP,

>> 1. Is there any book that I can refer? I feel I'm lost with so much of info/topics. Also with no audio for first class, it seems I don't have way to revisit the fundamentals discussed.

I don't use any one text book for DC. The material is collected and collated from multiple sources. However, wikipedia is your friend in case you need more detail on particular topics. Also pls check the early blogposts for yourbatch on analytics-yogi.blogspot.in where some additional links and material was putup.

>> 2. It will be extremely helpful if you can please provide some practice papers and solutions. (hope that's possible)

The exam is open book-open notes. The questions are all short answer quetions (no essay length stuff) for more grade-ability and objectivity. I can;t make any promises regarding the practice exam at this point as I plan to modify the exams I have from previously for this batch as well.

>> 3. Could you please also clarify whether Facbook assignment is group or individual H.W.?

Individual. Because each of you has to pullup your own FB data.

Hope that clarifies.

Sudhir

Ciao.

Friday, November 14, 2014

Make-up Assignment

Hi,

Make-up Assignment in lieu oif survey filling:

Pls watch this ~ 20 minute video carefully. It features Scott McDonald of Condé Nast holding fort on where MKTR is headed.

“Social Technological and Economic forces affecting Marketing Research over the next decade”

Now, for your HW, pls answer a few simple Qs (True-False, fill in the blanks variety) about the above talk in the following survey:

Questions for Make-up Homework.

HW Notes:

(i) This is an individual-only HW. Since it involves no R, consulting peers is not permitted.

(ii) I found that using earphones works great in making out what the speaker is saying much more clearly than ordinary speakers. FYI.

(iii) Deadline: The HW should be completed and submitted latest by midnight 10-December.

Any Qs etc, pls feel free to email me or use the comments section below.

Sudhir Voleti

Thursday, November 13, 2014

Interesting links from different facets of the DC course

Hi class,

Wide range of topics we'd seen in the DC course. Some of you asked for more sources and reading material. Pls find the same below (in no particular order) and totally optional only:

1. Recall the google glass example we'd seen in class? Well, here's a Gigaom article on the Future of the wearables market.

2. Recall the first example in the network analytics class on world international call patterns? Well, here's the associated Atlantic article on a World mapped by phone calls. It nicely illustrates how much visualization of networks can tell us.

3. More from the Atlantic on how its now technologically feasible to arrive at one's Identity. Big Data Can Guess Who You Are Based on Your Zip Code

4. Recall the habit patterns class we'd covered? Here's an article from HBR blogs on How Customers Get Hooked on Products.

5. There's an undercurrent somewhere in the program that spells the words "data science". This link here offers a rounded perspective on what precisely is data science. This follow-on link here describes 8 concrete steps you must take to become a data scientist. Yes, R features there. Apt read for all CBA students, IMO.

-------------------------------------------------

These links below are more technical in nature. And are even more optional reading than the ones above. I'd suggest revisiting the below links after a couple of more terms are done in the program.

6. This will be kinda boring to many perhaps. But here's an Academic journal paper on Behavior prediction using social networks

7. And here is an excellent set of slides for computing basic metrics in network data from r-bloggers.com. BTW, you should consider subscribing to their newsletter, if you are into R.

8. More R here. An excellent intro to general R and then some network basics along with code and examples workshop style.

That's it for now. Will update as more comes in.

Ciao.

Sudhir

Session 4 Homework for CBA Batch 3

Class,

Individual homework:

Fill up this survey below (on perceptions of what constitutes IT capabilities in a firm). If you have any issues with doing so, let me know and I will assign alternate individual homework.

IT capabilities survey



Group HW:

1. Pick up any well-known brand- product or service. E.g. Xbox360 or Jabong or iphone6 or Nike.

2. Collect 3 sets of data for it:

  • (a) 100+ consumer reviews from either flipkart or Amazon India
  • (b) 500+ tweets
  • (c) 50+ articles from Googlenews or any other news aggregator sites.

3. Feel free to either use R or any other means you know of to collect the data (e.g. Python, chrome scraper etc.). But clearly mention the data collection tool used.

4. For each set of data, perform the following analyses:

  • (a) General wordcloud using both TF and TFIDF weighing schemes. Update stopwords list to filter out noisy or irrelevant terms.
  • (b) Sentiment analysis. Display wordclouds separately for the top 50 most positive and most negative words.
  • (c) Identify the top few most positive and most negative documents. Read them and speculate on why they are so positive or negative about it.

5. Session 4 HW submission format:

  • Use a plain white blank PPT.
  • On the title slide, write your group name and the names + ISB students IDs of all group members.
  • Give your homework an informative title (include name of the product/brand you chose).
  • Have 3 sections in your PPT - one corresponding to one data source and separated by separator slides.
  • As slide separators, mention the source of the data. E.g., "Data source: Amazon Consumer reviews" or "Data Source:Twitter" and so on.
  • For slide headers, use format "TF Wordcloud" or "Positive wordcloud" and so on.
  • Save the slide deck as session4HW_yourgroup.ppt.
  • Put all the raw data you collected, the code you used and your PPT in a zip folder (so that I can replicate your analysis if need arises). Save the folder as session4HW_yourgroup.zip and upload in in the dropbox on LMS before the deadline.

Any Qs etc., let Atreyee or me know. Feel free to use the comments section to this post for any Q&A or discussions.

Sudhir

Session 2 Group Homework for CBA Batch 3

Hi all,

This homework covers sessions 1,2 and 3, i.e. problem formulation, construct assessment through qualitative research, and questionnaire design for primary data collection.

Group HW:

Consider the following Business problem.

A firm is planning to build a smartphone app that offers location-based services.

The app will collect details about deals, discounts etc from stores on one side and lets inform subscribers about these deals when they are within one cell tower range (roughly a km) of the business establishments where these deals are being offered.

The firm is targeting people below age 35 in the middle and upper-middle class in metropolitan India. The firm however wants to know what about the target segment's app usage habits in general. What types of apps do people use? Why? How many have used apps to transact business online (e.g., pay for orders placed) etc?

Your tasks will be to (1). conduct some exploratory/ qualitative research to find out what constructs underlie people's app based propensities and behaviors. Think of running a small focus group, or conducting a few in-depth interviews with knowledgeable people. (2). Formulate the problem in terms of a D.P. and a few R.O.s that correspond to it. (3). Design a questionnaire centered around measuring the constructs of interest you have identified. Read parts 1,2 and 3 in the HW below.



HW Part 1: Problem Formulation

  • Q.1.1. Write a decision problem (D.P.) to describe the business problem of interest.
  • Q.1.2. Write a few research objectives (R.O.s) to address this D.P. (pls use format specified for R.O. in class slides).

HW Part 2: Construct Analysis

  • Q.2.1. Conduct some exploratory/ qualitative research to find out what constructs underlie people's app based propensities and behaviors. FOr example, you could run a small focus group discussion, or conduct a few in-depth interviews with knowledgeable people.
  • Q.2.2. List a few major constructs you find from your data collected in Q.2.1. that are of business interest.
  • Q.2.3. Pick any one construct you have listed in Q.2.2. and break it down into a few aspects. Ask yourself what motives, means and opportunities drive the behavior associated with construct.
  • Q.2.4. Make a table with 2 columns. In the first column, write the names of the aspects you came up with. In the second column, corresponding to each aspect, write a Likert statement that you might use in a Survey Questionnaire to measure that aspect.

HW Part 3: Web-Survey Programming

  • Q.3.1. Build a web survey using any free online websurvey tool of your choice. E.g., surveymonkey.com or zoomerang.com offer free websurvey services.
  • Alternately, you can try Qualtrics, the ISB subscribed survey software. Instructions for how to setup a qualtrics account using your ISB email have been uploaded on LMS
  • Ensure your questionnaire is "complete" i.e. has an introduction, a section for the psychographic Likerts, a demographic section, and some gateway questions and SKIP logic.


Session 2 HW submission format:

  • Use a plain white blank PPT.
  • On the title slide, write your group name and the names + ISB students IDs of all group members.
  • Give your homework an informative title.
  • For slide headers, use format "HW Part 1: [Slide content description]" and so on.
  • Pls mention clearly the Question numbers you are solving in the slide body. Use fresh slides for each new article
  • Use a blank slide to separate HW Part 2 from HW Part 1.
  • Provide a working link for your websurvey on a fresh slide titled "HW part 3".
  • It is advisable to run a pre-test. Perhaps take the survey a few times to ensure clarity, readability, working SKIP logic etc is in place.
  • Save the slide deck as session2HW_yourgropup.ppt and put in in the dropbox on LMS before the deadline.


HW submission deadline: midnight of 10-December-2015. That's it from me. Any Qs etc., let Atreyee or me know. Feel free to use the comments section to this post for any Q&A or discussions.

Sudhir

Tuesday, November 11, 2014

Session 4 Classwork files on LMS

Hi all,

Sorry about the delay in updating the blog and the LMS.

Pls find on LMS R code, data and instructions for sessions 4 (text analysis).

That for session 5 (network analysis) will come in the next couple of days.

As CBA students, my expectation is that you will:

(i) diligently follow the instructions given,

(ii) read and understand the R code line-by-line before running it,

(iii) run the code and replicate the classwork examples,

(iv) discuss any issues etc that arise here on this blog by using the comments section,

(v) solve the group homeworks by tweaking and customizing the R code as required, and

(vi) provide constructive feedback where possible.

Instructions:

1. Unzip contents of the zip folder

2. Open Rstudio. File menu --> Open File --> textanalysis R code.R

3. the textanalysis R code.R file will open as an additional window (on the top left) in Rstudio)

4. To run any lines, select them and click the Run icon on the top right of the window. Ensure internet is connected.

5. Read the lines before running as some require input from your side (which files to read in etc)

6. The zip folder contents are self-contained and hopefully should run smoothly. However, if you encounter issues, pls let us know.

7. Pls email aashish_pandey@isb.edu with a copy to Atryee in case of any R related issues. Your group homework for this session will be up soon, in a few days. Pls ensure you are comfortable with this code before the homework arrives.

Thanks.

Sudhir