6 ms·
ArXiv receives $10M for upgrades
- armchairhacker 3y agoarXiv is a great resource. At least in my field (Programming Languages, but I imagine it applies to most if not all of Computer Science) it seems like most papers are on arXiv. I see no reason to not make papers free. Most things cost money because either they're material themselves, or cost money/material to produce which needs to be somehow recouped. But papers cost almost $0 to distribute and are created with grants which get funded by the results and discoveries themselves.
- omgJustTest 3y agoMy philosophy is that if people want to spend (waste) time on closed journals more power to them. Potentially when fields were small and niche this aggregation of topics might be useful. Today in modern search and computerized tools, the question is: why do we need them? There's nothing in academia that I feel is more of a racket than this scheme. I feel that _everything_ should be published and that natural discussion of said topic will drive the importance.
- CSSer 3y agoWould/could you apply this same argument to books?
- omgJustTest 3y agoI feel this is already true of books, the question is will a publisher put it out... books cost money to produce and by comparision to hosting are massive investments.
- jacobr1 3y agoI wish more self-published books would find a way to hire a copyeditor.
- EVa5I7bHFq9mnYK 3y agogpt to the rescue?
- jacobr1 3y agoActually, yeah. I think it could make a difference. Usually there is some basic spell-check, grammar check stuff already going on, but the kinds of things that would be easily fixed/identified by LLMs could be: - continuity checking - often moving sections of text around introduces temporal errors - style cleanup - quasi-duplicated text that came from multiple edits or start/stops being merged - fixing vague and unclear passages But probably what they won't do well in a near terms is provide more critical feedback, such as when to throw away large sections of book, or remove subplots, or where to dive deeper on other topics. I'm sure they can generate general suggestions for this kind of thing though.
- IanCal 3y agoWhat people usually picture for books is things written for profit by the author. That's very different from academic papers. If you're talking about monographs/etc then that's a different beast.
- jltsiren 3y agoToday, with modern search and computerized tools, one of the primary heuristics I use for determining which preprints to read is "do I know any of the authors". If I don't know you, your ideas are most likely not worth reading, even if they appear to be about a topic I'm interested in. There are so many papers published every year and so little time to read them properly, and the search tools that exist are not that good. If you don't have an established reputation, your best bet is submitting your work to a relevant journal or conference, where the peer reviewers will usually give you a fair chance. Google Scholar is a good example of how bad the tools are. Its recommendations are apparently based on things like the papers you publish, the papers you cite, and the papers that cite your work, which sounds superficially reasonable. That worked well enough when I was doing my PhD within the boundaries of an established academic field. But when I started doing more interdisciplinary work, the recommendations quickly turned into garbage. A better recommendation system would understand that there is a field I work in, an upstream field I take ideas from, a downstream field I contribute to, and several fields further downstream that use the work I contributed to.
- generationP 3y ago> migrate to the cloud Oh great, looking forward to a captcha every time I download a paper. (FWIW, arXiv is one place I trust not to make such stupid decisions, but migrating to a cloud often means the decisions are no longer made in-house...)
- mistrial9 3y agoI suspect this cloud move is more for security -- to keep others out -- than to improve the product. As you mention, many decisions will no longer belong to Cornell. This is indicative of larger geo-political trends.
- twosdai 3y agoIt's possible that they jump straight to the hybrid solution that many companies now make. Like migrating to the cloud could be just to have backups of papers in s3 with failures to download them initially then pass through to the s3 location seamlessly. It could also be trying to rewrite everything to run on lambda. My point being that migrating to the cloud means different things to different companies/groups.
- generationP 3y agoThat'd be good.
- jxf 3y ago> arXiv was founded in 1991 by then-Los Alamos National Laboratory physicist Paul Ginsparg, Ph.D. ’81, prior to his return to Cornell in 2001. I had no idea arXiv was that old. I honestly only heard of it and started using it less than 10 years ago.
- beezle 3y agoI still find myself typing in xxx.lanl.gov It used to redirect but doesn't anymore :(
- waveBidder 3y agothe most risqué place to store your preprints
- PaulHoule 3y agoIt was an email listserv before it was a web site.
- jessriedel 3y agoDifferent disciplines began regularly posting their papers to the arXiv at different times. The order was (very) roughly High-energy theory ~1992 Other physics theory ~1994 Physics experiment ~1996 Math ~2000 CS ~2008 Some disciplines flipped nearly overnight, while others took several years of slow growth before posting to the arXiv become the default.
- choppaface 3y agoWhy is this eating up NSF money / taxpayer money?! Megacaps like Google and Facebook are reaping tons of value from arXiv, i.e. (1) easy access to non-industry peer attention through a respected pre-print publisher, (2) major incentive to their employees and key component to performance review cycle, ... Google itself has out-published most universities in conferences like NeurIPS for the past few years. For Google alone to match the $10m would be a drop in the bucket-- a $10m grant would be an order of magnitude less than the compensation they're paying to their own employees who make heavy use of arXiv. NSF funds are scarce, precious, and it's immensely hard for lawmakers to divert funding to NSF versus e.g. defense. One thing arXiv needs is high-quality bandwidth and Google / Facebook have tons of that. Taxpayers are already paying tons for Google's several anti-trust trials, no need waste NSF budget for something industry can very very easily solve.
- perihelions 3y agoSeems a reasonable use of NSF money to me, to promote sciences by creating a common infrastructure for organizing the world's library of preprints. It's infrastructure of a sorts, and civil government is the ideal party to fund civil infrastructure. (Isn't it?) Not random FAANG's, and I don't see why they'd want to anyway—what they'd get out of it. FAANG's aren't charities; anything they do that looks philanthropic is some kind of trap.
- choppaface 3y agoI’m not saying it isn’t a bad use of NSF funds, I’m saying that especially in the AI space that industry use of arXiv is large enough that taxpayers deserve to have industry foot the bill.
- lkbm 3y agoI favor the "company pays tax, tax funds science infrastructure", cycle, not the "company funds science infrastructure" cycle. It reduces the risk of industry asserting influence over the infrastructure. Whether companies pay enough tax is a separate discussion, but they paid some taxes, and those taxes are now funding ArXiv. This seems like a good setup.
- tpmx 3y ago$10M seems way too much - if it's per year. Call me paranoid, but that kind of money invites the ESG crowd.
- tomrod 3y agoESG crowd? This would fund a team of 10 for four years + operating costs for active development for commodity cloud targets. This isn't a ton of cash.
- quickthrower2 3y agoOr 10x that if you get to choose another country :-)
- dheera 3y agoEh, inflation has happened, and taxes, medical bills, and real estate are insane these days. So much so that you can't live a "normal" life of the standards of 20 years ago with less than $500K+ in salary. (At normal "middle class" modest levels of life hat's $250K after taxes, $200K after medical expenses, $150K after rent, $100K after food, $50K after transportation, $25K after food ...) Sounds outrageous but most people have been brainwashed into thinking living with roommates in a cramped flat full of mildew, giving up all dreams of owning a home, and driving an old, unsafe car is the new norm. Given that, $10M isn't really that much, considering you need to hire software devs to do any of the work they propose, and those software devs are getting $500K+ offers everywhere else.
- tpmx 3y agoThis needs to be turned into a tshirt print: You can't live a "normal" life of the standards of 20 years ago with less than $500K+ in salary.
- Aaronstotle 3y agoI found it truly incredulous that you are trying to defend the position that $500k a year salary is middle class.
- yoshuaw 3y agoOh that’s great to hear! I really like arXiv and am glad it’s getting more funding! I hope this will enable arXiv to fix some of the long-time issues it has, including the way identity and attribution are handled. For example: the current system assumes people’s names do not change [1], which in particular negatively affects trans authors. This seems like it could be fixed if they had the bandwidth to, and I hope they will. [1]: https://www.kalzumeus.com/2010/06/17/falsehoods-programmers-believe-about-names/ https://www.kalzumeus.com/2010/06/17/falsehoods-programmers-...
- bigyikes 3y agoIt seems arXiv does support name changes, just not in citations and references[1] This seems… fine? To modify citations would be a strange alteration of history and an intrusion on authorship. If someone quotes this comment and cites it as “bigyikes” and later Dang changes my name, should he also edit the comments that quote me? I don’t think so, but maybe I’m in a minority here. [1]: https://blog.arxiv.org/2021/03/11/update-name-change-policy/ https://blog.arxiv.org/2021/03/11/update-name-change-policy/
- deleted 3y ago[deleted]
- yoshuaw 3y agoThe link you shared seems great, thank you! I was talking to a friend just yesterday who regularly publishes to arXiv, and my understanding was that changing author names on papers where they weren’t the primary author wasn’t possible. But from the link you’ve shared it seems they’ve made the steps needed to address that. I think this means I might be able to deliver some good news!
- solusn 3y agoIt much more negatively affects women who choose to or are compelled by societal expectations to take their husband's surname.
- 3y ago
- slowhadoken 3y agoI’m 100% on board with open source science.
- Osmium 3y agoI am really glad arXiv is getting more funding. It is an essential resource. For me personally, it’ll be really interesting to see if the frontend changes. No doubt there are some important improvements that can be made (a website can always be made more accessible, moderation tools, support for name changes as mentioned, etc.) However, to first approximation, arXiV’s website already seems almost like a platonic ideal. It reminds me of Craigslist. Simple HTML, loads fast, has the information and features you need but otherwise gets out of your way. I love it. The arXiv team deserves a lot of credit for what they’ve done to get it this far. It’s difficult to overstate how useful and transformative preprint servers have been to science.
- WanderPanda 3y agoLoads fast, except for the PDFs which load quite slowly
- open592 3y agoHave to download the file, then load it on the client. Not sure you can make it much faster.
- choppaface 3y agoIf the pdfs were served by some of the megacaps who posts tons of papers (e.g. Google or Facebook) then it would be an order of magnitude faster. And said megacaps would end up spending peanuts relative to the value they get from arXiv.
- codetrotter 3y agoIt’s better that ArXiv stays not relying on Google, Fb or any other tech giant. What I would like to see instead is first-party support for BitTorrent and magnet links.
- kawhah 3y agoAn order of magnitude faster sounds unimportant. Is all the time spent by researchers sitting around waiting for 0.4s for a paper to download really an issue?
- wilg 3y agoMake the papers web-native!
- generationP 3y agoThey seem to be trying: > In addition, arXiv will provide substantially better access to the visually impaired by producing HTML as well as PDF versions of its content. But it's a different question whether they will succeed. A compile-to-HTML pipeline needs to work with every single LaTeX package that people use, or at least work with most ones and fail gracefully for the rest. That was hard enough for compile-to-PDF, requiring the arXiv to keep every year's version of TeX Live permanently maintained.
- jahewson 3y agoIt’s a fools errand. pdfTeX didn’t even emit spaces (as in char 32) between words until a few years ago.
- quickthrower2 3y agoOCR?
- generationP 3y agoThough that one is perhaps to blame on pdfTeX rather than on the TeX format :)
- qumpis 3y ago"Ar5iv" does a pretty good job from my experience
- generationP 3y agoHuh, not half bad! Reporting a few bugs.
- elashri 3y agoSome people are angry that arXiv which is probably the most important repository of science are getting public funding because big tech companies are using it to publish their work, and they should fund that from their deep buckets. Are you aware of the consequences of what you are suggesting? If you are funding something, you have some authority over what you are funding. Directly or even indirectly (If you don't follow our suggestions, we might revisit our financial contribution). So you want to put the place that should be open and independent under the thump of corporations that does not always have the same goals as the benefit of the science and public interest?
- gerdesj 3y agoOpen Source is ... open source. There is nothing wrong with chucking money at Archive to help with costs, provided that money does not come with menaces. The basic idea is that if you profit in some way from an open repository of ... something then why not contribute to it? If enough people/orgs contribute then we are all good. Otherwise, if the org behind the project is an individual doing it as a hobby and they can't afford to continue then they might need support. In the end Archive and co have costs attached. You don't have to contribute but it might be nice if you did.
- dchftcs 3y agoPublic money indirectly comes from the most profitable companies, so it's not as bad as you suggest. I'd also say there are worse freeriders of academic research in the economy than big tech.
- elashri 3y ago> I'd also say there are worse freeriders of academic research in the economy than big tech Are you talking about Elsevier, et al. How dare you?
- fpgaminer 3y ago> Public money indirectly comes from the most profitable companies (In the U.S.) Corporations only account for 6% of federal tax revenue (https://taxfoundation.org/data/all/federal/us-tax-revenue-by-tax-type-2023/ https://taxfoundation.org/data/all/federal/us-tax-revenue-by...). And probably the most profitable companies are the ones that are the best at skirting corporate tax law...
- ftyers 3y agoHope it stretches to installing a TeX distribution that properly supports Unicode ...
- noqc 3y agoPrediction: Arxiv has no need of this money, as it has no need of upgrades. This money will exclusively attract salesman, all of whom will offer to make arxiv worse, for money. One of them might succeed
- gustavus 3y agoI concur. Seriously it's job is to act as a repository of PDF files. It does that it does that well. It doesn't not need a fancy new frontend that changes again in 18 months because "circles are in" it doesn't need to move to a "distributed load balanced fault tolerant web scale platform solution" (that costs about 2x in maintenance and engineering time as it currently does). It doesn't need to do anything of these things. It needs to serve documents and thats it. If there are problems with load they can always throw it behind a CDN since all of their contentt is largely static it fits nicely into the use case CDNs were developed to solve. Please don't change things.
- deleted 3y ago[deleted]
- ajuntapall 3y agoInternet movie pirates have been using bittorrent for 20 years. I think there is fraud here.
- chpatrick 3y agoWhy do you need $10M to host some PDFs?
- quickthrower2 3y agoBeing HN someone needed to make that comment. All tech companies can be reduced to “serving some bits and bytes”
- elbear 3y agoThe answer would still be interesting for those for whom it's not obvious. I include myself among those people.
- tamcap 3y ago* overhead - people cost a lot of money: this includes ongoing maintenance of the tech stack, including rewrites of obsolete parts, but also ops cost - preprinting can include more than one human touchpoint * converting PDFs to HTML is an annoying problem * searchability of the repo is likely an annoying problem * any new features that stakeholders want added (commenting, annotations, etc) * ongoing hosting / CDN cost
- chpatrick 3y agoHow much do each of those cost and how does it add up to 10M? If you get two overqualified people to work on it full time and pay them FAANG salaries it'll still be enough for decades. I can't imagine the hosting is expensive when 99% of the papers are a few megs at most. I'm just a bit confused because this is a site that already works really well and isn't technically difficult.
- chpatrick 3y agoDon't you find it just a bit mysterious how a site that already works perfectly well and does a technically pretty basic task can spend that kind of money?
- dilawar 3y agoGreat. BioRxiv comverts pdf to html. I hope arxiv does that too.
- jfdi 3y agoGood call. Surprised it’s not more given its value as (among other things) AI corpus.
- Incipient 3y agoIs moving something that has a seriously tight budget, to the cloud, necessarily a good idea?
- sideshowb 3y agoI wonder if they could use some of it to make an experimental post-publication review system. I'm not sure what the future of peer review is, but I doubt it would hurt for them to try.
- drones 3y agoI think it's antithetical to the philosophy of arxiv in the first place. It's a preprint service. Adding a review process would make it... a journal. Arxiv being impartial to the content it publishes is one of its' greatest strengths.
- sideshowb 3y agoAt a higher level of abstraction, arxiv is an attempt to fix problems with the review process. Attempting to fix more problems with the review process would.... Etc
- nojvek 3y agoArxiv is a blessing, I would not be where I am without Arxiv and all the folks openly sharing their research. I don’t like that deepmind gates their research behind nature and other publications. I acknowledge it puts them in the cool kids club but it feels unauthentic. Putting research out in the open moves humanity forward. Arxiv can say it made a little dent in the universe and I’m glad it’s getting the funding it needs to be sustainable. I just hope they don’t burn it out trying to scale needlessly.
- throwaway71271 3y agoI hope they just use the money for more bandwidth and maybe spend it on free api/download access so people can get big chunks of it easily and build all kinds of things (from alert on keyword to just stream of text to finetune models) And I hope they don't touch the frontend at all :)
- jacobgorm 3y agoPerhaps that means they can set the PDF title tag :)
- barrenko 3y agoThis will raise the entire world's GDP by at least 2%.
- zelly 3y ago$10M to host a listing of pdfs. lol let me do it.
- devilsAdv0cate 3y ago[dead]