Website Review
What is Open Source Research?
Open Source Research is the personal blog of a PhD student at Berkeley, used to share working notes on programming, data science and learning. Its tagline—"My daily sufferings as a PhD student at Berkeley"—sets the tone: informal, opinionated, and written from the middle of the work rather than after it.
Typical posts are practical and text-heavy. Examples visible on the site include "15 Principles for Data Scientists," notes on learning C++11, takeaways from a Stanford talk on learning to learn, a Mandelbrot fractal written in Python, and a critique of how Coursera and Udacity congratulate themselves. The data-science post is a numbered list of personal rules: be honest with data, build and share tools, keep studying graduate-level math and statistics, learn one language deeply and others well enough to communicate, and present your work publicly even when criticism is uncomfortable.
Who it suits
- A graduate student or early-career data scientist who wants a peer's raw notes rather than polished tutorials.
- Someone learning a language or tool who wants to see how another person structures their own study.
- Readers who don't mind profanity, strong opinions, and advice stated as absolutes.
Trade-offs
The value is in the directness and the range of topics; the cost is that posts are personal essays, not maintained documentation. Advice is tied to one person's experience at one point in time, so treat claims like "learn R" or judgments about specific tools as opinions to weigh, not settled guidance. Expect little structure, few updates, and no editorial filter.
If you want to use it well, pick the post closest to your current problem—say, the data-science principles list—and turn two or three of its rules into a weekly habit, then compare how they hold up against your own experience. For broader, continuously maintained material, pair it with a community resource such as Stack Overflow or GitHub.
What are the 15 principles for data scientists?
The 15 principles are a personal working creed for data scientists, published on a Berkeley PhD student's blog. The post introduces them as the rules the author follows in daily work; only the first nine are shown in the available page text, so the remaining six are not visible here.
The principles visible in the post
- Do not lie with data and do not bullshit. Be honest and frank about empirical evidence, and above all do not lie to yourself with data.
- Build everlasting tools and share them. Spend part of each day building tools that make someone else's life easier; humans are tool builders.
- Educate yourself continuously. Read graduate-level math and statistics, learn fundamentals rather than hallway explanations, read recent papers, attend conferences, publish and review.
- Sharpen your skills. Learn one language well, learn others well enough to communicate, treat SQL as essential, learn a compiled language, an interpreted language and R, and learn Unix tools such as sed and grep.
- Kick ass and amaze people. Do one thing every day that serves this purpose.
- Challenge yourself by presenting your work. Do not fear critics.
- Be generous with knowledge and ask questions. Do not hoard what you know.
- Develop your own ideas first, then listen to others. Use domain knowledge without being restricted by it.
- Hang out with people and talk to them. Learn how you can be useful in their projects and how their work can benefit yours.
How to use this
Treat the list as a discussion prompt rather than a standard. A new data scientist could pick one item per week — for example, "build and share one small tool" — and review it at the end of the week. A team lead could use principles 1, 6 and 7 as meeting norms: state uncertainty honestly, present unfinished work, and ask questions without penalty.
The list is opinionated and informal, with strong language and blunt judgments about tools and methods. Read it for the underlying habits — honesty, tool-building, continuous learning, communication — and decide for yourself where the specifics fit your stack and workplace. If you want the complete set, the original post is the place to look: Open Source Research.
How can I learn C++11 effectively?
Learning C++11 effectively means treating it as a modern language rather than "C with classes." The biggest practical shift is to use the standard library and move semantics from day one, not after you've already learned the old style.
A workable study path
- Get a modern reference. A book or course that explicitly covers C++11 (auto, range-for, lambdas, smart pointers, move semantics) will save you from unlearning outdated habits. Avoid pre-2011 material as your primary source.
- Write small programs daily. Compile with a C++11 flag (for example,
-std=c++11on GCC/Clang) so the compiler enforces the standard you're targeting. - Learn the standard library alongside the syntax.
std::vector,std::string,std::unique_ptr,std::shared_ptr, and the<algorithm>header do more for real productivity than memorizing language corners. - Read other people's code. Open-source C++11 projects show idiomatic use of lambdas,
auto, and move semantics in context. The blog Open Source Research documents one PhD student's notes on C++11 and related data-science tooling, which can complement a structured course. - Build one non-trivial project. A small tool, parser, or simulation forces you to combine classes, templates, and the standard library in a way exercises can't.
What to prioritize
| Topic | Why it matters | Learn it |
|---|---|---|
auto and range-for |
Cuts boilerplate, reduces type errors | Early |
| Smart pointers | Replaces manual new/delete |
Early |
| Lambdas | Everywhere in modern APIs and algorithms | Early |
| Move semantics / rvalue refs | Performance and correct resource handling | After basics |
| Templates and metaprogramming | Powerful but easy to overuse | Later |
A concrete next step
Pick one small program you already understand — say, a CSV reader or a word counter — and rewrite it in C++11 using std::vector, std::string, a lambda in std::sort, and std::unique_ptr where ownership matters. Compile it with the C++11 flag, fix every warning, then read the standard-library documentation for each facility you touched. That single exercise teaches more than several chapters of passive reading.
If you want a structured course, compare a few options on whether they explicitly target C++11 and include graded programming assignments, since passive video-watching alone rarely builds fluency.
What are the key takeaways from the 'Learning to learn' talk?
The page itself doesn't reproduce the talk's content. Its "Learning to learn" post is titled "My notes from the 'Learning to learn' talk by Stanford's Benjamin Von Roy", and the excerpt supplied here covers a different post — the author's "15 Principles for Data Scientists." So the honest takeaway is that this site is a personal PhD blog where the talk notes sit alongside posts on C++11, Python fractals, data-science principles, and criticism of Coursera and Udacity; the talk notes would need to be read on the post page itself.
What the surrounding material does show is the author's own learning philosophy, and it's a reasonable proxy for why a "learning to learn" talk would appeal to them:
- Fundamentals over shortcuts. The advice is to read graduate-level math and statistics rather than accept a hallway explanation of a method.
- Depth plus breadth in tools. Learn one language well, others well enough to collaborate, and treat Unix,
sed, andgrepas everyday skills. - Build and share. Spend part of each day making tools that make someone else's work easier.
- Teach and be questioned. Present your work, ask questions, and share knowledge rather than hoarding it.
- Form your own view first. Develop ideas before soliciting others' domain insights.
These are the author's stated principles, not a summary of Von Roy's talk — treat them as one Berkeley PhD student's framing rather than the speaker's actual claims.
Next step: open the "Learning to learn" post directly on Open Source Research to read the notes themselves. If you want the speaker's own material, look for Benjamin Von Roy's Stanford talk page or slides, since a note-taker's summary reflects what one listener found worth writing down.
How to judge the notes when you read them: check whether they give you something actionable — a study technique, a scheduling habit, a way to test your own understanding — or only inspirational phrasing. Notes that name specific methods and when they fail are worth keeping; notes that only say "learn deeply" are not.
How do I create a Mandelbrot fractal in Python?
Use a short Python program built on NumPy and Matplotlib: iterate the map z → z² + c over a grid of complex numbers, count how many steps each point stays bounded, and render those counts as an image. The page itself includes a post titled "A Mandelbrot Fractal in Python," so it is a reasonable starting point for seeing one working approach: Open Source Research.
H3 Steps that work in practice
- Choose a window in the complex plane, for example real parts from -2.5 to 1 and imaginary parts from -1.25 to 1.25.
- Build a grid of complex values with
numpy.linspaceandnumpy.meshgrid. - Keep an array
zof the same shape, initially zeros, plus an iteration-count array. - Repeatedly apply
z = z*z + conly where|z|is still below 2 (or 4 for the squared test), incrementing counts. - Plot the counts with
matplotlib.pyplot.imshow, usingextentso the axes show real and imaginary coordinates. - Optionally map counts through a colormap or a simple palette so the boundary structure is visible.
H3 A concrete reader scenario
A student wants one figure for a report and does not care about maximum speed. Start with a 1000×1000 grid and 100 iterations. If it takes too long, cut to 500×500 and 50 iterations; if the edges look coarse, raise iterations before raising resolution. If you later need animation or zooming, move the inner loop to a compiled extension or use vectorised NumPy operations rather than pure Python loops.
H3 Trade-offs to decide early
| Choice | Benefit | Cost |
|---|---|---|
| NumPy vectorisation | Simple, fast enough for still images | Memory grows with grid size |
| Pure Python loops | Easiest to read and debug | Slow for large grids |
| High iteration cap | Sharper detail near the boundary | Longer runtime |
| Smooth coloring | Avoids banding | Slightly more code |
If you want a broader data-science context around tool-building and Python practice, the site's "15 Principles for Data Scientists" post is the relevant companion: Open Source Research.
Next step: write the grid and iteration loop first, print the shape and a few counts, then add plotting. That separates a maths bug from a rendering bug.
What are the main criticisms of Coursera and Udacity mentioned on this site?
The site's main criticism of Coursera and Udacity appears in a post titled "Dear Coursera and Udacity! Don't congratulate yourself too much." The author—a PhD student writing about their daily work and study habits—objects to the platforms' self-congratulatory tone and argues that they overstate their own importance. The implied criticism is that MOOCs are not as revolutionary or as universally beneficial as their marketing suggests, and that celebrating themselves too much is unwarranted.
That is the extent of the criticism visible in the page evidence: the heading itself carries the argument, while the excerpt shows the author's broader style of blunt, opinionated commentary on learning and data science rather than a detailed critique of either platform.
If you want the full argument, open the post directly at Open Source Research and read the comments as well—the author's reply to readers often carries more of the substance than the headline. As a practical check, compare the post's date with the platforms' current course catalogs; a critique written years ago may target a different product than the one you would sign up for today.
User reviews (0)