Showing posts with label guest post. Show all posts
Showing posts with label guest post. Show all posts

Tuesday, 6 March 2012

The Unfulfilled Promise of Aggregating Institutional Repository Content (Guest Post)

Our thanks to Neil Stewart, Digital Repository Manager at City University London for the following guest post which raises some interesting questions for us all.
----
A very good question was posed on Stephen Curry’s blog by Björn Brembs recently (Curry and Brembs are a couple of the more prominent figures supporting the Elsevier boycott):
I’ve always wondered why the institutional repositories aren’t working with, e.g. PubMed etc. to make sure a link to their version is displayed with the search results. I mean, how difficult can this be?
This got me thinking, how difficult can it be? Aggregating and re-using institutional repository (IR) content at subject level is, after all, one of the promises of the Green road to Open Access.
The infrastructure is already in place, in the form of the many OAI-PMH compliant institutional repositories out there, and there is also the SWORD client, which allows flexible transfer of repository content. Some examples do exist- for example the Economists Online service, which harvests material from selected economics research-intensive universities, then makes it available via a portal. But (to my knowledge) there has been no work done to provide a way of e.g. ensuring all a repository’s eligible physics content is automatically uploaded to ArXiv, or all biomedical research to UKPMC.
Subject repositories have gained critical mass in certain disciplines (to add to the examples above, see also RePEC for economics, SSRN for social sciences and DBLP for computer science), meaning that if a paper doesn’t appear there, it’s far less visible. This means that the incentive to post locally is greatly reduced- yes, your paper will appear in Google, but a paper in ArXiv will appear both in Google and in the native interface of the repository where everyone else in your discipline is depositing.
So if the infrastructure is there and the rationale to create these links exists, why has it not been happening to any meaningful extent already? I suspect it’s because of the fact that the IR landscape is, by its very nature, a fragmented one. Those with responsibility for IRs (managers, IT people, and senior management) are understandably concerned with local issues: ensuring that IRs are properly managed and integrated with the university’s systems, as well as the usual open access and service awareness-raising and advocacy. Having time to think about the automatic population of ArXiv with papers from your home repository is probably pretty far down one’s to-do list.
That’s not to say that repository managers are oblivious to these issues- but here another problem arises. Few individual repository managers, I would guess, would think that they individually could negotiate with and persuade ArXiv that  automatic harvesting of physics content from their repository, and their repository alone, would be worth ArXiv’s while. This is, perhaps, where UKCoRR (or other national bodies- JISC perhaps?) might come in. If ArXiv or similar subject repositories could be persuaded of the merits of harvesting IR content (whether full text or metadata pointing back to IR holdings), it would allow all repositories to plug in to this system, and offer it as a service to academics (two for one deposit- local IR and ArXiv at the same time!)
So, what do people think? Is there any appetite for turning this into a project that UKCoRR members could take forward, perhaps with UKCoRR and/ or JISC oversight? Comments please!

Thursday, 22 September 2011

LSE Library and a REF call-out: lessons

In the second of his two guest posts, Neil Stewart identifies lessons from managing a call-out for publications for REF assessment at the London School of Economics.

If you would like to contribute a guest-post to the UKCoRR blog, please contact a member of the committee.
___________________________________________________________

At the recent RSP event on Readiness for REF (which I blogged about on this blog), my former LSE colleague Dave Puplett and I presented on our experience of managing a call-out for publications for REF assessment. The presentation was qualitatively different from the other presentations at the event, because it was on the managerial and organisational issues surrounding management of publications for REF purposes.

The slides from the presentation, which detail LSE Library's experience (full disclosure: I have now moved on to City University London, where I manage City Research Online) of managing the call-out and subsequent influx of publications, can be downloaded from the RSP site (Powerpoint link). Instead of re-hashing the whole presentation, I thought I would take the opportunity to detail some of the lessons learnt, which are hopefully of general applicability to repository people.

Lesson 1: dealing with REF matters puts you at the heart of things
When LSE Research Online (LSERO) was chosen as the de facto method of managing REF data, LSERO became a much higher strategic priority for the School. This is of course excellent for the service, and it had been the case that the LSERO team and Library management had been plugging away to make this happen for a very long time. However, it's also an opportunity that must be seized, because missing it could have meant that the repository would have been side-lined, and new methods to manage this process would have been found. At LSE, this meant that resources to adequately manage things had to be found, which meant diverting resources from elsewhere to allow this to happen. The REF is too important to ignore: get it right, by allocating adequate resource and managerial effort, and the repository gains profile and prestige; but getting it wrong could be disastrous.

Lesson 2: if you didn't talk to the Research Office before, you soon will
The REF call-out at LSE was instigated by the Research Office. While that team had been close allies during the RAE in 2008, the call-out meant that we really had to start working with them more closely well in advance of the REF. This soon fostered a productive relationship, and allowed us to use the Research Office's channels of communications with which to talk to departments. It also gave us the authority to standardise the way in which publications data was reported upon, since the combined weight of the LSERO team and Research Office left departments with little choice!

Lesson 3: it's possible to use the ePrints (and presumably DSpace) back-end to perform database query magic
If you're lucky enough to have a friendly IT person who can run SQL database queries (or if you have that skill yourself) then get in touch with them when you have to start thinking about REF matters. Being able to access then manipulate data direct from the repository's database is invaluable, because it allows you to create customised reporting data based upon any criteria you might wish to include.

Lesson 4: issues of disambiguation get thrown into sharp relief
Dealing with REF publications data brought up those old librarianship questions which are probably familiar to all of us. Two in particular came into relief particularly strongly:
  • Which department do academics really live in? Academics can have multiple allegiances, to their department(s), research centre(s) and other parts of the university (e.g. the senior management team). Where, for REF purposes, should an academic be placed? If "units of assessment" do not correlate with departments, how does the repository map this? These are of course as much questions for the Research Office as they are for repository teams, but nevertheless they must be tackled.
  • When do academics start (and finish) their careers with parent institutions? How much data from before (and after) these dates should the repository hold, for REF purposes?
The above points (and I'm sure other people can think of others) points to the need to have a CERIF-ied repository system, which links into other university-wide systems, and which may be able to solve these problems of ambiguity.

Lesson 5: in-press and submitted publications are hard to deal with
Be very careful about how forthcoming publications are dealt with. The problem here is one of recording this data in a non-public forum, which can still be reported back to departments in a useful fashion. Many academics will feel that you are jeopardising their chances of publication by including a citation to an in press item in the live repository without their say-so.

Lesson 6: don't let Open Access be forgotten about!
All of the above sounds like work that could usefully be done by a CRIS, and makes no mention of the primary goal of (most) repositories, which is providing openly accessible research. There is a great danger, in my view, that open access can be overwhelmed by the needs of REF reporting, particularly if the repository team has to devote extra resource to dealing with this. How to balance open access and REF is an open question, and one that the LSERO team are still pondering. One benefit of the REF exercise is that it has made LSERO "complete" (regarding citations, at least), which might be a way of further pushing the open access agenda from a position of strength.

I'm sure there are plenty of other lessons that people could add to this list, judging by discussions at this event and elsewhere- please add them (or any other comments) in the comments section below.

Friday, 9 September 2011

RSP Readiness for REF (R4R) workshop, 5th September 2011


In the first of two guest posts, Neil Stewart reflects on the RSP Readiness for REF workshop.

If you would like to contribute a guest-post to the UKCoRR blog, please contact a member of the committee.
___________________________________________________________

My name is Neil Stewart, and I'm the repository manager for the newly minted City Research Online repository, at City University London. I normally blog at City Open Access, if you want to keep an eye on developments of a repository which is still on a project footing, rather than a fully-fledged service.

The reason you find me writing here is because I was recently the recipient of two invitations: to present at the RSP Readiness for REF (R4R) workshop, held on Monday 5th September in London, and to blog about that event here at UKCORR's blog. I was happy to take up both invitations. What follows summarises some thoughts about the workshop, and on the Research Excellence Framework (REF) as it relates to repositories in general. It should be noted that the opinions below are my own, and do not necessarily reflect those of UKCORR. I'll be writing another post soon, which outlines the R4R workshop presentation which I delivered with my former LSE colleague Dave Puplett.

The event's emphasis (apart from the presentation which Dave & I delivered) can fairly be characterised as macro-level, since it discussed REF data submission in very general terms, and with specific reference to the CERIF metadata standard. CERIF is a flexible, extensible model for metadata about research-producing institutions, including (but not limited to) publications data. It is possible, using the CERIF schema, to model an institution's structure, then show how researchers and research outputs relate to that structure. This has obvious benefits for an exercise like the REF, which will (amongst other things) require submission of data on "REF-able" (i.e. high quality) publications at a department (or at least department-like "unit of assessment") level.

The morning sessions dealt with laying out the details of how CERIF and CERIF-compliant repositories could assist with REF submissions. The first session was an overview of the JISC-sponsored Readiness4REF (R4R) Project, and was delivered by Richard Gartner of Kings College London. R4R, just in the process of closing, has looked at the way in which repositories and other data management tools can provide CERIF-compliant data for REF submission purposes. Second was Keith Jeffery from euroCRIS, who gave the bigger picture as to how CERIF was developed and what it can do. This was followed by a panel discussion on the Measuring Impact Under CERIF (MICE) project, which is attempting to build in research impact data into the CERIF schema, and hence make it readily submissible for REF.

After Dave & my presentation and lunch, there were demonstrations of R4R plug-ins for the three major repository software types (Fedora, DSpace and ePrints). As an ePrints user, I was interested to see a demonstration of ePrints v. 3.3, which is to be released in the next few weeks, and contains some "CRIS-like" functionality. This is a kind of CERIF-lite approach, by which it is easy to create associations between researchers, research grants, research centres etc. to express CERIF-like linkages between them. These linkages can then be exposed in useful ways using ePrints web pages, but also exported as CERIF data to be re-used in other systems, or for a REF submission. This seems to me an interesting development, and one we may have to look at here at City.

The final session featured the obligatory break-out groups. I was assigned to a group which discussed the question: "Do you think CERIF is now a more viable option for your institution to use for its REF submission?" A variety of subjects were covered, as ever with these type of discussions. The two main points I took from it were the fact that CERIF provides the opportunity to provide an "open" citation model by modelling linkages (including positive and negative citations) between publications, outside of the "walled gardens" provided by Scopus and Web of Science; and that, for CERIF to work within my institution, there is the somewhat intractable problem of knowing to whom to speak to find out if, for example, the HR database can be plugged into the repository to transfer CERIF-formatted data between the two systems.

All in all, an interesting and timely event. Keep an eye out for my post on LSE Library's experiences of conducting a mini-REF, coming soon!